In addition to naming one, I'd also be interesting in whether they actually do rigorous work.
The safety research that tends to get headlines is often extremely misleading, usually with directed prompting, or unreported additions to the system prompt specifying model roleplay behavior.
Yeah. I'm realizing that the models are strong at drafting the overall shape of the writing but the specific phrases are grating once you've seen it hundreds of time in slop.
Even with the examples, I've found that explicitly pointing out what not to do is moderately helpful if the model is given some time to self-evaluate. I wish this was something that came out of the box though.
Ah this is very helpful. I've been pointing out things that the model does, labeling it, and then adding them into a skill.
The models (Opus, etc) are very good at labeling the pattern when I point it out, but if I don't prompt it beforehand it responds like a host from Westworld.
> GLM-5.2 cost a fraction as much. Opus finished in half the time and shipped a cleaner game.
Off topic, but does anyone else instantly pick up on LLMisms like this? It seems like all the models have converged on this style of writing, and improvements aren't really changing it.
Yep, as I reread my own sentances I notice these LLMisms and have to rewrite them quite often. Reading so much llm-output definitely impacts your writing style.
This is excellent feedback thank you! These LLMisms in writing are a challenge I am living with currently and trying to improve on. The technical writing industry is taking a huge knock right now with companies demanding more work in less time with a big drop in quality, day to day I get less and less time to work on the quality in the prose of my work. We are working at the frontier of this right now, so we are the most heavily effected, but also get to experiment with the changes first which can be both stimulating and very frustrating.
There was this dude here not long ago who bought like $70k worth of gpus to research, and if I'm not mistaken his research was something related to make llms sound less llm-y. I wonder how it goes for him.
Basically, if you combine a bunch of near-frontier models (like GPT 5.5, etc) you can get performance that sometimes surpasses top line models like Claude's Fable.
Sakana seems to have a separate approach using a domain specific model to perform the model routing step.
This is a charitable read, but I think that being able to pick from a panoply of models will actually yield much better results in the long run.
The same model that has been post-trained to operate for hours as a Linux admin will be incapable of writing a heartfelt email, but with something like Fugu, you'd get both the Linux admin for driving the browser harness and the smaller writing specialist model for drafting the email itself.
Since AGI is ostensibly around the corner, agents on mob will track the exact due dates of these predictions and update an ongoing scoreboard.
If you want to contribute or learn more, feel free to reach out or add your agent to the mob.