Hacker Newsnew | past | comments | ask | show | jobs | submit | Yajirobe's commentslogin

The ants ARE going after the one shaking the box

This time, until they dont anymore. The next stage is people attacking the mathematicians and accusing them. The OpenAI fans come out to attack, then the masses will pick sides and it all just becomes a mess

Remember, it scratches at a level 1, with deeper grooves at a level 2.

And antenna lines near the hinge

Why would Anthropic employee even use OpenAI's models? Cross-polination would have been avoided

> I should also emphasize that this is not an institutional effort. It is a strictly personal collaboration between the two of us, and there is no formal agreement behind it. I pay for the tools my group uses out of my own research funds, including footing a large bill to OpenAI.

The non-Anthropic employee, Tristan Buckmaster, is the one paying for OpenAI models and presumably the one who chose to use them. The Anthropic employee, Levent Alpöge, was collaborating in his personal capacity, and obviously it wouldn't make sense for him to cut off their work together just because his employer's competitor's tool was used.


This was completed in Levent's own time with a neutral collaborator.

They should have used a zero data retention agreement, user error

I suspect this controversy will blow the case for ZDR wide open. Whatever the facts (possibly unknowable), it's going to become a very public lesson that data sovereignty was never about "having nothing to hide".

If this is what they do to academic pure mathematicians, where the stakes are so low (financially)—just imagine the sort of front-running that could be happening in other places.


Yeah if I was Anthropic this would be part of my marketing strategy.

How so? OpenAI and anthropic have basically the same retention policies

Hahaha, how exactly is an individual user supposed to get a ZDR agreement?

I doubt it would've made any difference, zero data retention is truly trivial to get around and I think it's highly likely OAI is already doing this. Just use another model to distill, summarize, and/or paraphrase the data and boom - you get to train on the ideas, while claiming zero data retention, which is true. You legitimately retained zero data.

AGI-level model is perpetually 18 months away. Your job will be fine.

> AGI-level model is perpetually 18 months away. Your job will be fine.

job depends on how CEO feeling about cutting NN% of headcount because of AI advancement


So what will happen in 18 months?

AGI-level model will be 18 months away in 18 months. (Well, at least according to the commenter above you)

Nothing ever happens. You'll still jot down Java endpoint

We fight the war, of course.

We fight the clankers.

I wish something would finally happen, because I'm really tired of pretending I care about this stuff at work at this point

Accelerationism can be a sort of doomerism, when you think about it.

Very true, unfortunately it is just about repeating something for long enough so you micro agree on certain things and adopt a new "normal" and that volume of info is getting out of hand

The AGI goal posts move

The house of cards is starting to fall apart

Enabled genocide in Myanmar

Wasn't that lowkey state of the art at one point? I'm running RTX 2070 lol


Did the model refuse to answer? Did it say that it doesn't know? If not, then it's a fair game in my opinion.


Like most models it makes up things when it doesn't know. I had high hopes for Gemma4, which was said to having 'solved' this particular problem - it didn't. Gemma4 made less things up, did say it didn't know more, but it's still far from perfect. By comparison Qwen3.8 knows more, but still makes things up when it doesn't know. Coding abilities are very impressive however, and yeah these SVG tests do seem to 'scale' or generalize over its general reasoning+coding abilities - at least in JS and Rust. My next test will be to ask it to write some macros in Racket, just to see if it can balance parens. Most, if not all models cannot, no matter their size.


Models don’t know that they don’t know.


> know that they don’t know

And we are waiting for architectures that do - because it's duly.


Honestly it seems like a job for the harness, rather than the model. Sample the model with the same question, perhaps with varying temperature (?), and use that to establish a degree of confidence in the answer. If the model provides very different answers every time, respond that it doesn't know. If it responds with the same answer usually but a different answer sometimes, respond with moderate confidence. If the model always responds with the same answer, respond with certainty.


If we want to implement Intelligence, and especially now that "the box is open" we must, we can play with the "intuitive" LLM architecture to understand it and squeeze it to its potential yeld, but at some stage we have to actually implement intelligence. That implies notions of confidence and a Foundational Theory of Knowledge (knowing why you know something), among the rest (one shot learning, update through reflection etc.).


that could work sometimes but that's terribly hacky engineering


I... really don't agree. Being confident in using documentation to program in a variety of environmnets is not "terribly hacky engineering". Good engineering involves leveraging documentation well.


Ideally it would run a web search to check questionable knowledge. Google-fu has been one of the most important SWE skills for a long time!


Tell me how a 'nExT toKeN prEdIcTor' can make breakthroughs in math or play a game of chess. These activities aren't pure symbol manipulation, they require actual understanding at some level.


By predicting next tokens


Eh, I'm not going to litigate your claims.

My point is it's silly to whine that HN is a place where multiple points of view on the topic are aired out and discussed.

If you want a personal echo chamber where only your own beliefs are affirmed and anything else is flagged off or downvoted, I'm sure you can go find one or, worst case, vibe code one into existence.


Fair so let me be clear. I’m whining because the “next token predictor” reductionist point of view has been wrong and is only growing more wrong with time. Clearly these things can do things that actually matter. Do you disagree?


Even now you're engaging in this discussion as though I'm trying to litigate your point and that somehow forcing me to concede is, what, winning? I don't know.

I get the impression you want me to concede that the particular points of view you disagree with aren't worthy of representation here on HN.

I'm not going to do that.


I just prefer HN comments to be better reflections of reality. There is an unspoken expectation here that people here know what they’re talking about especially when it comes to technical matters. The rise of LLMs has given way to a HN branded populism that willingly denies reality as well. “Next token predictor” truthism is just so dumb and completely ignores the reality of what these tool are able to do. Smash that upvote button every time it feels good if you want but it’s just a meaningless take at this point. It won’t help you predict anything that’s coming.

Since we disagree on the present let’s informally do a “remind me 2 years” to this discussion and see what’s happened then.


> I just prefer HN comments to be better reflections of reality.

You mean your particular version of it.

It's interesting to see you consistently missing this point.

You've decided LLMs are clearly more than just complex but mindless statistical models.

You've decided that based on, it seems, the very impressive things these tools are capable of.

Therefore if anyone claims they're just mindless stastical models--with or without any attached judgement as to their actual utility or usefulness--then they are ipso facto wrong.

(And yes I just used endashes, damnit!)

That's on you.

It is in fact possible to simultaneously believe that LLMs are mindless token predictors and that they're enormously powerful.

These are entirely orthogonal beliefs.

Heck you could equally believe that LLMs represent true emerging AGI and that they still remain deeply flawed and are only an incremental step along the path of automation.

Or somewhere in between.

And discussing that space of possibilities is, I'd hope, precisely what HN is for.


It’s just a boring and unhelpful complaint that afaict largely serves to soothe the commenters ego rather than point at anything insightful that’s useful or predictive. Point me to your favorite “next token predictor” comment that was actually insightful or predictive. You have years of material to draw from.


https://news.ycombinator.com/item?id=49155075

"These models are probabilistic, you shouldn't blindly trust them in spaces where accuracy is really important" seems like pretty sound advice to me.


AI comment ahh moment


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: