Hacker Newsnew | past | comments | ask | show | jobs | submit | CJefferson's commentslogin

Wow, that is widely disingenuous, I don't really think there is any excuse for that, I don't believe someone deep in compression algorithms wouldn't know they could adjust the block size, and 512GB is a huge block size for bzip3, as it needs to basically all be in memory so you can't pretend that's just 'the standard value'.

> 512GB is a huge block size for bzip3

Sorry! That was a typo, it should have been 512MB (now fixed). Still huge.


What’s interesting is it’s not obvious how this is leveraged to ‘cure disease’. But I’d love to know.the advantage of this is there is a clear measure of success. Here is a rule language. Prove this. You are done when your proof passes. You can sit quietly and spin for billions of tokens.

How does that work for drugs? We can’t let AIs make millions of test drugs and try them out on people.


I don’t attach a raw dump of all the conversations I have with colleagues to my commit messages, and wouldn’t want to work in a place where I did.

Sometimes I tell Claude ‘don’t trust file X, it was written by a colleague who writes awful tests’. I don’t want that in my commit messages.


I worry this doesn’t check correctness. I’ve been finding Claude is lately awful at folllowing instructions, I’ll ask it to implement the algorithm from a paper and it will do something simpler and slower and when challenged do it’s stupid apology thing. It can’t be trusted with anything I’d put in a paper, it lies too much. 4.6 couldn’t do as complex tasks, but it would do what was asked.

Benchmark does check correctness. Team is working to write a paper that will likely contain failure mode analysis; checking for instruction following could be a good idea.

Agents are given instructions in markdown format, allowed to read data and libraries sandboxed in a Docker container, and evaluated on deterministic pytests on the outcome. Two things the team aimed to enforce to add tasks we could trust: - Scientific workflows are often simulations that are correct upto numerical tolerances (the scientist decides what's reasonable), so task verifiers' evaluate results of agent-written code within the tolerances. - There could be multiple solution codes to a scientific workflow, and the team tried to ensure the verifier tests accommodate those. Not overfit to the oracle reference code, written by the scientist.

Instruction following is implicitly assumed, if the model gives up and doesn't complete the task it counts as a failure because the verifier tests fail.


> I worry this doesn’t check correctness

Then it's not a valid benchmark. I agree though they're not reliable enough to just put results in a paper.


This is completely different to my experience and my reading of the research. Please link me to any of this research you are talking about.


Meta-analysis, very large, shows small correlation between unhappiness and social media use: https://jamanetwork.com/journals/jamapediatrics/fullarticle/...

Large AU observational study. Moderate social media use correlates to the happiest kids, high use and no use correlates to unhappiness: https://jamanetwork.com/journals/jamapediatrics/article-abst...

Swedish obervational study, after adjusting the data there was no correlation: https://www.jahonline.org/article/S1054-139X(26)00162-X/full...

I can't find the specific inverventional study I mentioned. As I recall, it was mixed but possibly I remembered wrongly.

Some studies do show a significant correlation between unhappiness and social media use, but causation has not been established and the bigger, better done studies have weaker correlations. People believe in causation because it suits their bias against social media.

I dislike social media but I doubt it's harmful. This substack summarizes my view on what really makes children unhappy: https://substack.com/home/post/p-186087964 (school). I'd guess that lack of freedom is what makes school so awful; modern schools are like little totalitarian states, where every minute of your day is accounted for and there is no privacy.


“There are studies that do not line up with my opinion but they aren’t as good as the studies that do” lol


This is a pretty unimpeachable causal identification strategy finding Facebook's rollout caused worse mental health issues https://www.aeaweb.org/articles?id=10.1257/aer.20211218

It's hard to do similarly high quality studies today because no one rolls social media out gradually anymore.


This is why I’m trying to move to open Chinese models — because I will be able to use them forever, while the older Claude models which I genuinely enjoyed writing short stories with have now been deleted, replaced with hypothetically cleverer models which produce text everyone hates.


I also kind of miss how "unhinged" the earlier models were.


What about censorship?

> I will be able to use them forever

Where will you run them when powerful enough GPU and RAM are only sold to hyperscalers?


Everyone censors for their core jurisdiction/audience. The Enlightened West just calls this guardrails


"nothing happens in 1989" is censorship. "no I won't tell you how to build a bioweapon for genocide" is guardrails.

idk about you but i WANT the second thing, because I like to be alive.


It's much more civilian than bioweapons.

You only need to mention Protected Group Of The Week (I'm one of them and I like to research and read about history, so that makes it extra challenging) or anything resembling negative human emotions (guilty on that front as well), and the model screeches to a halt.

Just because OAI doesn't want another headline like "Chatbot convinces teen to off himself"

I understand this in principle, but I'm not convinced that dulling everyone's knives is better than figuring out how to keep them out of kids' hands


> Just because OAI doesn't want another headline like "Chatbot convinces teen to off himself"

And so you consider that censorship and not guardrails? Huh...


That was not the censorship part.

"dulling everyone's knives" isn't either...

> "no I won't tell you how to build a bioweapon for genocide" is guardrails.

I call that censorship too. I'm curious enough that I want to know about such things. I don't want any limitations on what I'm allowed to understand and know about.


Your freedom ends where my begins. Welcome to study chemistry and learn things from first principles. but if you want to just ask for the practical steps of making a genocidal bioweapon then honestly I don't know you or whether you are honestly "merely curious" and so prefer your freedom ends there?


> Welcome to study chemistry and learn things from first principles.

As if your censorship was not going to kill that too. Can't even ask Fable about aminoacids without getting blocked. So much for "learning from first principles".

> but if you want to just ask for the practical steps of making a genocidal bioweapon

Nothing wrong with practical steps. Just because I know how to do something, doesn't mean I'm actually going to do it.


> As if your censorship was not going to kill that too.

I didn't say "with LLMs". Last time I checked they still teach chemistry at unis and schools.

And yeah, guardrails are not perfect. Honestly I don't think good enough guardrails are possible, it's all eventually defeated or becomes silly. And yes that should be one of the reasons the technology as a whole is banned.

But until then, guardrails are guardrails and not censorship in a pretty obvious way. If you refuse to see that, be my guest. I personally like to live.

> Nothing wrong with practical steps

No thanks from me


You make it sound like the DNC.


The Democratic National Committee?


US models censor and restrict more things than Chinese models by quite a margin.


So you choose bad over worse and pretend it's good


I haven't said anything about "good", I just pointed out that there isn't exactly an alternative to Chinese models if censorship and restrictions are regarded as bad for longevity. You can't really do better than the Chinese models for longevity; US models are by far the worst in this regard. So "What about censorship?" is an absolutely hilarious question to ask when Chinese models are presented as an alternative. Yes, what about it? They have less than the obvious alternative from US labs, and where is it you imagine you'll find less censorship?


Censorship or guardrails?

"nothing happens in 1989" is censorship. "I won't tell you how to build a bioweapon for genocide" is guardrails. I like the second one because I like to be alive.


I don't share your enthusiasm for the guardrails and I don't believe at all that they accomplish what they're supposedly created for.


not enthusiastic about them, more like very disturbed without


> Where will you run them when powerful enough GPU and RAM are only sold to hyperscalers?

Do you think that fabrication will never progress (in volume) than what we have now? The hyperscalers are already having trouble paying the bills, they can't keep this up forever.


The hyperscalers will be bailed (maybe not all of them but enough). US economy will crash if not. And whatever is made will go to them, because they pay more (thanks to US taxpayer bucks among other things) than any regular person. First they build on land then they build in space.


I dislike all censorship, but US models are much more censored, I often find myself using Chinese models to get answers I want.

Now, of course I’d prefer no censoring, but I live in the world we live in.

I’m working in the assumption that (like today) there will always be somehow on openrouter, or similar, who will host a model I want to run.


How do Western models censor?


I’m guessing they’re referring to things like refusals if it thinks your request may be related to building a bioweapon, or hack someone else, etc.


Much less than that. Anthropic won’t let its best model discuss my digestive system issues, or how to install apps on my own Peleton exercise bike.


Some of this seems public safety not censorship.

By censorship I mean things like "nothing happens in 1989". By public safety I mean "no I won't tell you how to build a bioweapon for genocide".


My big problem with leaving Google is I’ve clicked on far too many ‘login with google’ boxes. In retrospect this was a mistake as it now ties my own domain to google basically forever, but lots of sites don’t even have the option to undo it.


A lot of those sites will let you just reset password, send an email to the Google account's associated gmail, and then you're done. I suspect that it isn't thought through as an official feature in such cases, but it does often work.

Additionally, if you signed up with an email and password, you can often just sign in via SSO with a matching email address for convenience and still have the underlying password remain working. Again, possibly not an intentional feature (but rather something like using the email reported via SSO or via email verification as a primary key).


I switched away from Google and haven't had a single issue going from 'login with google' to a regular email/password FWIW. Every site has offered some kind of painless recovery.


This is a bit surprising to me because I switched emails some time ago, and most sites were painless, but there were a fairly long tail of sites that had no self-server mechanism to change email from one to another. A fair number of them required phoning in to read off my new email to a CSR rep. Another pile of them had no mechanism to change email at all, their only solution was to open a new account.


You can change the MX records of the domain to point at Fastmail (or some other provider) without losing the ability to log into associated "Google Workspace" accounts.


I thought this would be a blocker too but you can both host your email somewhere else and continue to login with google. You still need to keep the google workspace account enabled and active, you just point the MX record elsewhere.


Same; I was a sucker for GitHub SSO as well. I have no idea how to unwind it all, but (like Google SSO) it certainly had its benefits while I was using it. I remember there was an “open” SSO provider - OpenID maybe - back in the 2000s when a federated identity made interacting with blogs and such much easier. It’s the same with GitHub today, with various 3rd parties who need access to your code. They don’t have a “give me an git+ssh url” and so you almost need GitHub.


i’m not sure i’ll ever delete my google accounts, but my goal was to stop using them. if someday an email comes in, it’ll still get auto forwarded, and if i need to login, i’ll have access to it.


> and if i need to login, i’ll have access to it.

You hope you can ... I didn't login to a Google Account for a few months now and it suddenly wants a lot of personal details to do additional "verification" because I am "suspicious".


Did it also give you a tip to "log in to your usual Wifi network"? I have a feeling all large corps let their heuristics run wild with little to none supervison.


I don't even bother with forwarding mail from Gmail, it's just a black hole for marketing emails that I'd only occasionally glance at when I'm looking for a discount code.


Go here to see all of the apps linked to your account: https://myaccount.google.com/linkedapps

Then go through each service and email them to have them try and move your account to new accounts with a traditional username/password login. Some might even be able to change your login method to username/password without requiring a new account!


I was paying for Google Business for my family, and then decided to switch to iCloud, but left only my own account on Google, because of that reason. Also, some services requires you to log in with a Business Google account, regular gmail accounts don't work (ie: Granola).


Hmmm, that is an interesting pickle. Some are probably fairly easy to swap over but it is so dependent on each service.

Alas, short term convenience can lead to long term troubles.


Make a new account then?


I'm a software developer. This is clearly bad, and we can decide what it should be called. It might not be 'hacking', but you are clearly abusing the computer to steal a space in a class you shouldn't have.


My gym class operates on 'write on a piece of paper, cross off your name if you want to cancel'. I could cross someone else's name off and write mine in. They would have trouble figuring out it was me who did the malicious cross-off.

That wouldn't make it remotely acceptable of course.


I find it amazing people don’t believe this. I lost about 50kg and people treated me better. They didn’t randomly shout ‘fat fuck’ in the street for a start.


I can't believe there is even a discussion. The better you look, the better people treat you and being better looking also means being of a normal weight.

I literally have a weight number where if I am north of the number, I can tell people treat me differently and not as nice and if I am south of that number, people treat me a lot nicer (gasp, Women even hit on me). People are fooling themselves if they think their weight does not matter to the world around them.


Skinny fuck!

(This is me say good job, and congrats on your new found health :) )


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: