Hacker Newsnew | past | comments | ask | show | jobs | submit | galsapir's commentslogin

yep. i think it talked about a lot in the context of research, but not enough (or maybe im just not exposed to it) in relation to the arts.


thanks for reading it properly and engaging with the argument!

writing is hard, expressing ideas cleanly is harder! working on it.


curious where the disagreement lands: the claim i'm least sure of myself is that measurement alone already counts as activation (nothing in the weights changes, so it's a looser sense of the word than usual) the part i'd defend harder is the eval -> reward one: once a benchmark becomes the thing you train against, its flaws stop being measurement error and start being incentives. if you're pushing back somewhere in there, i'd genuinely like to hear it


really interesting that its basically almost 80% claude opus..


yeah its really counterintuitive i think; i.e, getting the right framework and structure for this to work probably isn't trivial, models really hate playing well together. i wonder how their version would fair in real world use.


i feel like i've had exactly the same thought in the past :-0 might even have written about it. feel your pain


As someone wise told me, it's just a procrastination ouroboros


jj


sometimes I also feel it tries to optimise for "per line coverage" over more "real, complex use cases" type tests


hey that's pretty cool. I think I still prefer "distill HN" cleanliness though. What made you create this.


I didn’t make this lol; just something cool I’ve found


axon discharge is brilliant. adopting.


oh sorry! didn't catch the one Thanks, I'll comment there


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: