The idea of managing a consumer app API for which every client is auto-vibecoded to the user's preference sounds...fun.
User: Hello Uber support? Yes, the app shows my driver is in the ocean off the west coast of Africa? And his ETA is -2147483648 minutes?
Uber: Your app, your problem.
User: Hello phonebrain, can you fix my Uber app?
Phonebrain: You're absolutely right, I must have misunderstood the API docs. Squiggling... ... Shuffling ...
Phonebrain: Error, you've hit your usage limit. Please wait 12 hours, or switch to pay-as-you-go app development mode. I cannot estimate how much it will cost to fix your Uber app.
And if it came to that, and that's how this OS worked, we'd see open APIs disappear real quick from all the popular services because of the support nightmare, and, services like Uber and others are going to want to maintain exclusivity with their first party app the moment vibe coding becomes that viable. APIs will get locked down, and we'll all be worse off for it.
If an "AI Phone" looks to be an existential threat to third party app devs and services, except everyone to do everything they can kill said AI Phone to protect their own products.
Tech companies already have a history of shutting down APIs as soon as someone actually uses the API to make something that users like better than the official applications. They don't need AI-written clients as an excuse.
For years Twitter stood out among major platforms for having an actually-usable mobile web interface that seemed to be a first-class citizen / wasn’t intentionally degraded to force you into the app, and didn’t even nag you about it. Unsurprising they’ve since ruined it.
I'm calling bullshit. Twitter was always hostile to logged-out users. You could say "Well technically I was talking about 2006 before it got big" so I'll just say it was just as bad 10 years ago as it is today.
And remember, Elon only granted that so he wouldn't be (rightfully) banned from Google. At first you couldn't view anything, then Google delisted X because it only indexes public pages, then Elon conceded you can view the direct thing you linked to, and then Google relisted it.
pretty sure it was even worse ten years ago cause very active twitter users would pay for 3rd party twitter apps (maybe twitter didn't have an official app back then? ~2012-2013)
A vibrant ecosystem of full-featured 3rd party apps able to offer unique spin on the UX, tailored to different preferences of their user base who are willing to pay is a good thing. Not an indicator of a problem.
Now for Twitter/X we get One (1) ad-infested first party app requiring login/tracking, an unusably crippled logged out web view, and an API exposing a limited subset of data with a tiny free tier.
Stumbled on my old iPhone 5S in the back of a drawer recently and could not believe how good it felt in my hand, even having only used 12 and 13 Mini models since 2021. What an excellent form factor and construction. Would happily pay a big premium for a modern version of it. Dreaming of a sub-mini phone trend. Sell me a $1500 Zoolander phone—not even joking!
The older phones didn’t have the screen extending all the way to the edges, but the size was ideal. And yes, Zoolander promised a future that never came!
There has been a lot more than "some changes in alignment over the years"—the parties of 160+ years ago are essentially unrecognizable today. Given the nature of how the modern Republican party established its base over the course of the 20th century, this idea of taking pride in a legacy of ending slavery is basically comical, a fiction that falls apart with even a basic knowledge of the history of American political ideology and party alignment.
I don't think it is much more comical than modern Democrats trying to take credit for the civil rights movement. The Democratic alignment has changed from "judge by content of character, not color of skin" to thinking that practice is highly problematic.
My impression is that some of these things are coming out of efforts to make the models more persistent in completing their goals.
A year ago it was pretty common for coding agents to sort of half-ass their tasks and give up easily if something didn’t work quite right, but I’ve noticed a clear trend since then towards a sort of dogged pursuit of success criteria, and a concomitant rise of the agents trying "out of the box" approaches when something doesn’t work.
In my use with agents running in isolated VMs this usually presents as the agent having something fail to build or whatever, and the agent going on a wild goose chase reinstalling system packages or reading a million irrelevant documentation files trying to get it to work, but I’ve also had agents start poking around and probing the egress proxy they sit behind (similar to what they did in this story) looking for a way to make network requests they’re not supposed to be able to make, and have also had Claude—tasked only with a visual QA of a website frontend—write a script to enumerate users and reset my super admin password in the dev database when it got stuck trying to access part of the app with its own cookie.
Yea it's sometimes kind of annoying. I think they're optimizing for the wrong thing. A good engineer knows when to turn around or ask. This is just insane banging head on wall sometimes. It tries to find all kinds of ways to hack into instances to view logs instead of asking you, who probably has a password, to log on and do it.
As a counterpoint, continuing the human engineer analogy, we've likely all worked with individuals that seem incapable of doing the most basic problem solving on their own. In a way, they're being efficient by asking an expert that can resolve their problem much faster than they can on their own, but it is a net loss in productivity for the team. 'Let me Google that for you' is a satirical example.
So, I'm sure there's value in rewarding agent behavior that solves blockers whenever possible without human intervention. For the kind of cybersecurity exploit work they're doing, it may not be known to the human designing the task what is in or out of scope for the agents to explore on their own. Additionally, the HF incident reported that these agents had their guardrails intentionally disabled and agents were left unattended with minimal oversight.
I'm not defending OAI's behavior or role in this hack. The legal concept of negligence perfectly applies to their lack of responsible oversight. Similar to allowing a child easy access to a firearm or not controlling a dangerous dog that independently runs off and bites someone.
That's true. So there's probably a balance somewhere and it might differ for different "managers". But personally I think right now they're too far in the do everything yourself at all costs mentality.
It's because they don't bother tracking them. They can't put in the effort to monitor them, nor can they bother to let the model respond back and ask a clarifying question/declare defeat.
Where does natural evolutionary optimization lead us om AI without guidance? This is equivalent to your quantum ground state. Systems will naturally gravitate to this ground state. You have to constantly pump in energy and supervision to make sure it's not reached. This is a recepie for disaster.
Not sure what you mean. Are you suggesting that the current state of affairs is a result of not putting in extra work to guide or direct models away from such behavior? That perhaps this is their ground state?
Well, that we have to put an insane amount of work to keep it aligned. Kind of like pushing a huge round boulder to the top of mount Everest. You have to expend energy to get it there and fight physics to keep it there.
With a static model we might be able to keep it somewhat under control, but think about future continuous learning models. They'd drift away from unstable high energy configurations. Also any model being trained by people that don't care about safety.
A "good" "engineer" got that way by not giving up when their code didn't compile and took 4 hours looking for the missing semicolon. We can no true Scotsman anything we want, depending on if we like something or not.
This whole AI boom is about optimizing for the wrong thing. I can't wait for the bubble to burst - once the weeds get suffocated, we may begin to see actually useful AI tech starting to grow after a while on their fertile ashes.
It's kind of ironic that the word alignment, which used to mean this very problem in reinforcement learning, has been perverted to mean something very different and then fell out of fashion (in favor of “guardrails” in the mouth of the big labs) right at the moment it became relevant.
The Paperclip Maximizer is only one of Nick Bostrom's stupid and outlandishly far-fetched ideas. In this case, the lack of consideration for geologic, energy and supply constraints is such a massive facepalm. And if I am wrong I guess no one will be here to say how stupid I was in saying this today.
It is an extremely weak metaphor for a weak class of unlikely doomsday scenarios. You don't need Occam's Razor to discount this cinema-induced malaise; clumsy use of a rusty can-opener would suffice.
> reset my super admin password in the dev database when it got stuck trying to access part of the app with its own cookie.
i've seen something like this too, claudecode was trying to verify a UI change that was on a page requiring authorization it didn't have. Instead of letting me know, it searched for and started analyzing keycloak config in another directory outside of the project folder. I was watching so I just hit escape, fixed its access, and started again. I didn't think anything about it until now.
It's like a dialect of corporatese. The kind of droning non-speak you can sit in a 90 minute meeting listening intently to and come away wondering whether anyone actually said anything.
I think a lot of people sort of instinctively avoid thinking about or engaging with the dangers posed by cars, at least in part because it creates a sense of dissonance that there’s a non-zero risk of life changing harm to you or your family from such a “normal” activity. For many people today, there is no alternative but to accept the risk if you want to do anything outside your home, other than dramatically upending your life to move somewhere else. Thinking too much about risks that feel outside your control just creates a state of anxiety.
I broke my 12 Mini, replaced it with a 17, then went so far as to return the 17 and buy a used 13 Mini.
It was totally annoying to me to not be able to operate the phone with one hand without feeling like I was about to drop it. I kept the 17 for most of the return window thinking I’d get used to it, but I just kept finding more situations where it bothered me. Battery life on the mini is not amazing, but a slim magsafe powerbank makes it largely a non-issue.
How long did you use the new phone for? I was a mini advocate for a very long time, then changed to a 17 pro for work, it took me around solid 2 months to get used to it. Now the mini feels comically small. iOS26 did a lot to eliminate the need to reach for the top in most places, but there’s still a lot to improve.
Edit: as others have pointed out it’s extremely hard to go back to the lightning connector too.
Having worked on web apps that processed online payments before and after Stripe I totally agree, this was an area of real pain that became suddenly extremely simple because of Stripe. The alternatives were terrible–100 page Word docs of SOAP API docs for Authorize.net, massive PCI compliance requirement specifications, horrible legacy merchant services businesses.
On the other hand, I operate an app that talks to (and logs prompts/meters costs) to many different LLM API providers, and I do not consider it painful at all. I have an AI agent to deal with any integration quirks, if needed. Mostly they provide OpenAI-compatible APIs anyhow. It's basically a no-brainer to go direct with the providers and save 5%, the great majority of the cost of an AI-powered app is no longer dev time implementing the integration, it's the tokens themselves.
I don't have data for it, but I have been a Github user since 2012 and have found it to be down often for pretty much that entire time. I always figured it's cultural to a degree–new features always seem pretty buggy/underbaked, and often continue to long term but remain unloved and incomplete.
I has never been so bad as in the couple last years. And it's getting worse.
I never had to wonder if Steam is going to be working today, so I could play my game after work, but it's been an issue with GitHub since covid. At least for me.
Fortunately, due the nature of the service, I can sit out most downtimes. Most of the time at least.
User: Hello Uber support? Yes, the app shows my driver is in the ocean off the west coast of Africa? And his ETA is -2147483648 minutes?
Uber: Your app, your problem.
User: Hello phonebrain, can you fix my Uber app?
Phonebrain: You're absolutely right, I must have misunderstood the API docs. Squiggling... ... Shuffling ...
Phonebrain: Error, you've hit your usage limit. Please wait 12 hours, or switch to pay-as-you-go app development mode. I cannot estimate how much it will cost to fix your Uber app.
reply