Is perceived model degradation after launch just regression to the mean?
Many model launches follow the same arc: amazement in week one, "it's been nerfed" a month (or week) later. I'm currently experiencing the same thing with Opus. But I feel I not only experience these things with AI: other stuff also gets harder after the first week sometimes. Is it just regression to the mean? Are we comparing launch-week highlights with everyday output, and getting disappointed that it's not so good as that one amazing new thing we did and got me on a high? Is it loss aversion strengthening that? The gains we start to expect, the losses we are hit by? Or do we start with our best use cases and simply run out of them? And then if feels like the model is underperforming, while it might be our part? Don't we try and thinker as hard as we did in the first week? Even with Opus 5.5, it took time and iteration to get certain things right? Or are we so expecting and used to constant progress, that even a temporary plateau (the same model) on a trajectory still rising across releases feels like a regression? I see all the incentives and pressures there are for companies to reduce performance. I'm sure they do that for some part in some cases. I just wondering if we could seperate the two. Could we compare how this feels on commercial APIs and chatbots? Could we do that with blind A/B testing? Like an Arena?