Perhaps what OP is noticing is that models are plateauing for real world use cases (even while doing even more impressive things when you have unlimited tokens to burn), and their training regime to try anything and everything until "the task is complete" to the point of throwing spaghetti at the wall to see what sticks
I think they try pushing the models in a direction where they just go on until somehing is finished end-to-end. So basically building "the loop" into the model itself. E.g. with GPT-5.5 the model would often do only part of a job and get back to me with a recommendation for next steps and I would either say "Yes, go on" or "No, instead do X". And the newer models never even get back to me, they just keep on churning, deciding about the direction on their own. Sometimes they go into the right direction, sometimes they drift into something unrelated or just go down some rabbit holes for hours. I liked GPT-5.5 better, and I built my own tools around it to still get them to finish tasks end-to-end which works fine.
I was thinking more about the cultural problems that lead to stories like this
https://globalnews.ca/news/11676795/tumbler-ridge-school-sho...
Perhaps what OP is noticing is that models are plateauing for real world use cases (even while doing even more impressive things when you have unlimited tokens to burn), and their training regime to try anything and everything until "the task is complete" to the point of throwing spaghetti at the wall to see what sticks
I think they try pushing the models in a direction where they just go on until somehing is finished end-to-end. So basically building "the loop" into the model itself. E.g. with GPT-5.5 the model would often do only part of a job and get back to me with a recommendation for next steps and I would either say "Yes, go on" or "No, instead do X". And the newer models never even get back to me, they just keep on churning, deciding about the direction on their own. Sometimes they go into the right direction, sometimes they drift into something unrelated or just go down some rabbit holes for hours. I liked GPT-5.5 better, and I built my own tools around it to still get them to finish tasks end-to-end which works fine.