I agree somewhat with the way the agents behave but feel the opposite reaction. With Fable, I get exhausted because it's always dumping out paragraphs of text that explain one approach but have some secret gotcha thrown out in the last two sentences. Then I have to pause and consider the caveat and if it matters and it happens every single time Fable responds and that constantly needing to make a decision that could radically change the approach gives me decision fatigue. I much prefer how much more decisive Astra can be.
I'm the opposite. Every time I've let Sol/Astra be decisive, I ended up with an overengineered mess.
I much prefer getting alerted when there's more than 1 approach to the problem and it's discovered mid-implementation.
I don't want to do the grunt work of writing code, but I do want to know the architecture and be responsible for the decisions.
Fable is also very good at pushing back when I propose something that will cost me. E.g. I'm working on a configuration layer above nix to manage my homelab fleet declaratively, and I tend to get into "config as new language", where Fable just goes - let's not do that.
Oh man! This also is a pet peeve of mine with Fable. I will look at what it's doing and say "Shouldn't it be done this way?" and then it will spend forever arguing with me that it should be done the way it wanted to do it. It seems to get stuck in a certain way of thinking and will insist its way is right until I can really prove it - or just go over to Astra.
That same story happens without LLM involvement. I’ve also seen LLM suggest updating a lib. So unfortunately nothing about this story speaks to a difference between LLM and human capabilities.
What if you have 10-20% certainty the AI being built will kill everyone? Go on a rampage and land yourself in prison just so Lab B or China can win the arms race and their AI can kill everyone? Probably not. Quit your job? Sure.
You don't have to go on a rampage in an American lab. If Chinese labs were ahead I'd bring about the same discussion.
Though separately if the United States (or China or anyone) believed one country or another was truly going to achieve something akin to a metaphorical AI Supremacy maybe you nuke them, or at least the labs/researchers. Many a sci-fi movie has been built on a similar "first strike" premise.
> In fact, it is now the identification of a promising problem which is the scarce and precious resource. We have now seen that even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential. The incentives may now be pointing in the direction of no longer sharing any promising research directions with the broader community, which would reverse centuries of traditions of open science and do serious long-term damage to the future of the field.
> it is now the identification of a promising problem which is the scarce and precious resource
This is by no means new. Perhaps it is even more extreme now. Literally my first 1:1 with my PhD adviser back then, he told me that the most important thing about a researcher is the quality of the problems he picks.
Shouldn’t this very capable model they’ve developed be able to identify promising problems? That’s what I’d expect from how the model is being presented and advertised.
That's more or less what they did according to their announcement. They fired it at a whole bunch of high end math problems and merely concentrated all efforts on one after it made some promising progress.
It seems like that’s the opposite of what happened. They started attacking the problem when they got a wind of a possible solution from certain individuals.
But they didn't know which problem had a possible solution. So they fired it at a huge amount of problems and then merely focused on the one that turned out to be promising. The model found the promising path itself, they just re-allocated the available resources once it became apparent.
But the previous administration denying a fair primary process gave us the present one due to forcing an unpopular candidate. They collude with each other so much that they're effectively one party.
I bounce between Sol high/medium and Luna max. I don't know why you'd use anything between Luna max and Sol medium. Luna is so extremely cheap and cranked up to max it does anything I'd want Terra to do for a fraction of the cost. What is Terra for?
One thing I've really noticed with Luna Max is its speed. I've got a review script setup on a custom Pi extension. Luna finds some issues/some false positives, while Sol finds issues but disregards false positives. The biggest thing is Sol finishes in about half the time.
reply