Wow, that is widely disingenuous, I don't really think there is any excuse for that, I don't believe someone deep in compression algorithms wouldn't know they could adjust the block size, and 512GB is a huge block size for bzip3, as it needs to basically all be in memory so you can't pretend that's just 'the standard value'.
What’s interesting is it’s not obvious how this is leveraged to ‘cure disease’. But I’d love to know.the advantage of this is there is a clear measure of success. Here is a rule language. Prove this. You are done when your proof passes. You can sit quietly and spin for billions of tokens.
How does that work for drugs? We can’t let AIs make millions of test drugs and try them out on people.
I worry this doesn’t check correctness. I’ve been finding Claude is lately awful at folllowing instructions, I’ll ask it to implement the algorithm from a paper and it will do something simpler and slower and when challenged do it’s stupid apology thing. It can’t be trusted with anything I’d put in a paper, it lies too much. 4.6 couldn’t do as complex tasks, but it would do what was asked.
Benchmark does check correctness. Team is working to write a paper that will likely contain failure mode analysis; checking for instruction following could be a good idea.
Agents are given instructions in markdown format, allowed to read data and libraries sandboxed in a Docker container, and evaluated on deterministic pytests on the outcome. Two things the team aimed to enforce to add tasks we could trust:
- Scientific workflows are often simulations that are correct upto numerical tolerances (the scientist decides what's reasonable), so task verifiers' evaluate results of agent-written code within the tolerances.
- There could be multiple solution codes to a scientific workflow, and the team tried to ensure the verifier tests accommodate those. Not overfit to the oracle reference code, written by the scientist.
Instruction following is implicitly assumed, if the model gives up and doesn't complete the task it counts as a failure because the verifier tests fail.
I can't find the specific inverventional study I mentioned. As I recall, it was mixed but possibly I remembered wrongly.
Some studies do show a significant correlation between unhappiness and social media use, but causation has not been established and the bigger, better done studies have weaker correlations. People believe in causation because it suits their bias against social media.
I dislike social media but I doubt it's harmful. This substack summarizes my view on what really makes children unhappy: https://substack.com/home/post/p-186087964 (school). I'd guess that lack of freedom is what makes school so awful; modern schools are like little totalitarian states, where every minute of your day is accounted for and there is no privacy.
This is why I’m trying to move to open Chinese models — because I will be able to use them forever, while the older Claude models which I genuinely enjoyed writing short stories with have now been deleted, replaced with hypothetically cleverer models which produce text everyone hates.
You only need to mention Protected Group Of The Week (I'm one of them and I like to research and read about history, so that makes it extra challenging) or anything resembling negative human emotions (guilty on that front as well), and the model screeches to a halt.
Just because OAI doesn't want another headline like "Chatbot convinces teen to off himself"
I understand this in principle, but I'm not convinced that dulling everyone's knives is better than figuring out how to keep them out of kids' hands
> "no I won't tell you how to build a bioweapon for genocide" is guardrails.
I call that censorship too. I'm curious enough that I want to know about such things. I don't want any limitations on what I'm allowed to understand and know about.
Your freedom ends where my begins. Welcome to study chemistry and learn things from first principles. but if you want to just ask for the practical steps of making a genocidal bioweapon then honestly I don't know you or whether you are honestly "merely curious" and so prefer your freedom ends there?
> Welcome to study chemistry and learn things from first principles.
As if your censorship was not going to kill that too. Can't even ask Fable about aminoacids without getting blocked. So much for "learning from first principles".
> but if you want to just ask for the practical steps of making a genocidal bioweapon
Nothing wrong with practical steps. Just because I know how to do something, doesn't mean I'm actually going to do it.
> As if your censorship was not going to kill that too.
I didn't say "with LLMs". Last time I checked they still teach chemistry at unis and schools.
And yeah, guardrails are not perfect. Honestly I don't think good enough guardrails are possible, it's all eventually defeated or becomes silly. And yes that should be one of the reasons the technology as a whole is banned.
But until then, guardrails are guardrails and not censorship in a pretty obvious way. If you refuse to see that, be my guest. I personally like to live.
I haven't said anything about "good", I just pointed out that there isn't exactly an alternative to Chinese models if censorship and restrictions are regarded as bad for longevity. You can't really do better than the Chinese models for longevity; US models are by far the worst in this regard. So "What about censorship?" is an absolutely hilarious question to ask when Chinese models are presented as an alternative. Yes, what about it? They have less than the obvious alternative from US labs, and where is it you imagine you'll find less censorship?
"nothing happens in 1989" is censorship. "I won't tell you how to build a bioweapon for genocide" is guardrails.
I like the second one because I like to be alive.
> Where will you run them when powerful enough GPU and RAM are only sold to hyperscalers?
Do you think that fabrication will never progress (in volume) than what we have now? The hyperscalers are already having trouble paying the bills, they can't keep this up forever.
The hyperscalers will be bailed (maybe not all of them but enough). US economy will crash if not. And whatever is made will go to them, because they pay more (thanks to US taxpayer bucks among other things) than any regular person. First they build on land then they build in space.
My big problem with leaving Google is I’ve clicked on far too many ‘login with google’ boxes. In retrospect this was a mistake as it now ties my own domain to google basically forever, but lots of sites don’t even have the option to undo it.
A lot of those sites will let you just reset password, send an email to the Google account's associated gmail, and then you're done. I suspect that it isn't thought through as an official feature in such cases, but it does often work.
Additionally, if you signed up with an email and password, you can often just sign in via SSO with a matching email address for convenience and still have the underlying password remain working. Again, possibly not an intentional feature (but rather something like using the email reported via SSO or via email verification as a primary key).
I switched away from Google and haven't had a single issue going from 'login with google' to a regular email/password FWIW. Every site has offered some kind of painless recovery.
This is a bit surprising to me because I switched emails some time ago, and most sites were painless, but there were a fairly long tail of sites that had no self-server mechanism to change email from one to another. A fair number of them required phoning in to read off my new email to a CSR rep. Another pile of them had no mechanism to change email at all, their only solution was to open a new account.
You can change the MX records of the domain to point at Fastmail (or some other provider) without losing the ability to log into associated "Google Workspace" accounts.
I thought this would be a blocker too but you can both host your email somewhere else and continue to login with google. You still need to keep the google workspace account enabled and active, you just point the MX record elsewhere.
Same; I was a sucker for GitHub SSO as well. I have no idea how to unwind it all, but (like Google SSO) it certainly had its benefits while I was using it. I remember there was an “open” SSO provider - OpenID maybe - back in the 2000s when a federated identity made interacting with blogs and such much easier. It’s the same with GitHub today, with various 3rd parties who need access to your code. They don’t have a “give me an git+ssh url” and so you almost need GitHub.
i’m not sure i’ll ever delete my google accounts, but my goal was to stop using them. if someday an email comes in, it’ll still get auto forwarded, and if i need to login, i’ll have access to it.
You hope you can ... I didn't login to a Google Account for a few months now and it suddenly wants a lot of personal details to do additional "verification" because I am "suspicious".
Did it also give you a tip to "log in to your usual Wifi network"? I have a feeling all large corps let their heuristics run wild with little to none supervison.
I don't even bother with forwarding mail from Gmail, it's just a black hole for marketing emails that I'd only occasionally glance at when I'm looking for a discount code.
Then go through each service and email them to have them try and move your account to new accounts with a traditional username/password login. Some might even be able to change your login method to username/password without requiring a new account!
I was paying for Google Business for my family, and then decided to switch to iCloud, but left only my own account on Google, because of that reason.
Also, some services requires you to log in with a Business Google account, regular gmail accounts don't work (ie: Granola).
I'm a software developer. This is clearly bad, and we can decide what it should be called. It might not be 'hacking', but you are clearly abusing the computer to steal a space in a class you shouldn't have.
My gym class operates on 'write on a piece of paper, cross off your name if you want to cancel'. I could cross someone else's name off and write mine in. They would have trouble figuring out it was me who did the malicious cross-off.
That wouldn't make it remotely acceptable of course.
I find it amazing people don’t believe this. I lost about 50kg and people treated me better. They didn’t randomly shout ‘fat fuck’ in the street for a start.
I can't believe there is even a discussion. The better you look, the better people treat you and being better looking also means being of a normal weight.
I literally have a weight number where if I am north of the number, I can tell people treat me differently and not as nice and if I am south of that number, people treat me a lot nicer (gasp, Women even hit on me). People are fooling themselves if they think their weight does not matter to the world around them.
reply