The thing that bothers me here is less that the model cheated and more that it found a way to improve the score that the people running the test didn't intend. That's a pretty nasty failure once you start giving these things more control
I must say I am becoming a huge fan of Deepseek. They keep putting out capable models and actually tell people a lot about how they built them. Even if I don't understand every part of it, I'd rather see companies show their work than give us a few benchmark charts and call it a day
What I find interesting is how AI has changed the cost of maintaining two native apps enough that a decision that made no sense a few years ago is worth revisiting now
The higher price seems less important if it actually gets the job done with fewer tokens. I'm still very worried that this will end up coming back to bite us, by becoming more expensive once they inevitably nerf it. Every major model provider does that now after all
The common thread between this and the other incident seems to be agents finding somewhere they can leave information for other agents. Once they discover a writable surface, it basically becomes shared memory for them
This is really cool, but $33 for a single generated world makes it hard to see this being useful for games just yet. Not to mention how this would go in a much larger project
There are free open source tools that convert osm data into 3D models, for example https://osm2world.org/. Also you could use freely available national 3d point cloud data, at least for some areas. both are $0 options. btw I am not saying $33 dollars here is expensive; just giving you some comparisons
The recent Sonnet models have been disappointing for me personally which is why I'm going look into using Opus/Fable as the planner and Flash as the executor. Let the expensive model handle the hard thinking and use Flash for implementation and tests so that I can stretch the Opus/Fable usage further
reply