To be fair Codex and Claude Code don't have a lot of secret sauce either. This is not to say that the harness doesn't matter at all or that you can't do any better, but whenever they benchmark it always turns out that you don't lose that much with even a very minimal agent. If a different harness makes better sense for your workflow I don't see why you wouldn't switch (or make your own if you have the resources to do so).
I'm also sticking to first-party agents, but that's more because I feel no compelling reason to switch and CLI agents fit well with my spartan technology choices. YMMV.
> The ARC-AGI-3 scorecard is extremely misleading (...)
True.
> Regardless, the result is still valid (...)
If you think the game is rigged, the virtuous thing to do is to point that out and refuse to partecipate; making up your own rules is something I just don't understand, especially since the rule-abiding result would still have been SOTA.
> in the sense of passing the most famous benchmark designed specifically to measure AGI progress
The benchmark does not measure AGI progress or progress towards superhuman intelligence, as explicitly stated by the creators.
On the AGI question: surely you realize this depends on how we define the term? For example, one of the definitions OpenAI originally gave is "capable of doing most economically valuable work", which almost certainly Astra, as impressive as it is, would fall short of. I'm not saying it's a good definition, but as far as I'm concerned it's as good as any. More importantly, I don't think that it would change much if we said yes or no. I'm only bothering to take a position if it amounts to something.
This can feel as "moving the goalposts", and to some extent it is, but if done honestly "moving the goalposts" is how you make progress. Had you asked me 10 years ago I would have said that anything that could hold a conversation like GPT-4 could would probably have been wildly superhuman at almost everything. It shouldn't be hard to find ways GPT-4 was lacking, though. We see new things, we reassess and try again: that's how it's supposed to work.
They now give out resets you can choose when to use, but with a "best before" date. With Anthropic you don't even know what model you'll have access to next week. A few months ago I wanted to try a subscription to Gemini but I couldn't figure out how to give them money. The other one thinks he's Mecha Hitler.
To an extent I get it, this is not like a normal SaaS where you just buy seats, but their commercial offerings, especially B2C ones, feel like complete amateur hour. If they were selling literally anything else I'd have ran for the hills a long time ago.
I agree with you contra GP that "barely speak English" wildly misrepresents LLM capabilities, but I think they don't write very well, and your characterization is also an overstatement. They often use stylistic elements that don't make sense in context: pointless metaphors, abuse of bullet points and tables when a simple paragraph would do, excessive parataxis, misplaced emphasis, dense technical prose even when unnecessary, occasionally weird word choices.
Those problems are much worse in languages other than English, by the way. I'm not a native speaker but I always interact with LLMs in English because their prose in Italian is terrible, it often reads as though the text had been written in English and then translated.
Obviously there's nothing wrong with each of those stylistic elements in abstract, exactly like there's nothing wrong with em-dashes or the word "absolutely", but I would argue that writing well means choosing styles that match the content; what LLMs appear to be doing instead is forcing the message into a default pattern, which we percieve as forced and therefore tend to dislike.
To be clear: their writing is fine. The primary goal of writing is to convey meaning, and LLMs have been perfectly adequate at that pretty much from day one. I can get over their quirks.
The bar I set was "significantly higher quality than the median literate native speaker".
The bar you're setting is expert use of metaphors, restrained use of bullets and tables and parataxis, proper emphasis, and avoiding dense technical prose.
I suspect you have a higher opinion of the median literate speaker than is warranted.
If your work is more focused on statistics or pure modeling, then I agree R wins hands down. The issue is that most projects have "unclean" parts where you have to gather data from multiple sources, use connectors for services, S3 buckets and whatnot; dealing with that mess is where Python really shines.
AI probably changes the equation to some extent, but I still believe I'd rather maintain a complicated data pipeline like that in Python rather than R.
Agreed. If the R community developed more data pipeline frameworks, following the "tidyverse" way of doing things, R would be my go to choice for all data related work.
> His polemic against determinants is poorly motivated, misguided, and distracting.
What polemic? Defining the determinant as the unique multilinear alternating form satisfying certain properties is very normal (and in fact the only way that really makes sense for both finite- and infinite-dimensional vector spaces). There are zero unusual things with this book imo.
Wikipedia says "The book has a pure, proof-heavy focus and is aimed at upper-division undergraduates who have been exposed to linear algebra in a prior course." [1], so it seems to be a different category of book?
That's actually ok. In virtually all Latin dictionaries verbs are listed in the first singular person of the present indicative (e.g. you would find "sum" and not "esse"). It's just showing you the dictionary entry.
It's not really very useful in this particular case, as you can't actually navigate the "dictionary" any other way than finding the base word form somewhere in the text you're looking at.
> When OpenAI posted about their 10 breakthroughs, I saw lots of career research mathematicians say things mostly along the lines of “I don’t understand any of this it’s way over my head”.
Huh?
Mathematics is an extremely wide subject, it's perfectly normal even for two professional mathematicians not to understand each other's work. Have you considered that maybe they just don't work in that area?
Are you implying OpenAI's paper (which was, by the way, edited by humans and provided Lean certificates for most of the proofs) is actually gibberish? That's flat-earth levels of conspiracy.
I'm also sticking to first-party agents, but that's more because I feel no compelling reason to switch and CLI agents fit well with my spartan technology choices. YMMV.
reply