That does actually kinda make sense, but it works very differently... the politeness is anything but superficial (even counters the racism and xenophobia to a limited degree, but also hides it, coming out to ±0), and the latter two are indeed present but coded as to be ashamed of rather than proud… (unfortunately the coding of shame is also different and doesn't spurn action).
It actually makes it harder to tackle the discrimination, since it's not openly admitted. Glass walls and ceiling.
I think raw brain energy is not a fair comparison. Humans are not willing and able to serve requests at identical competence all hours of the day. You have to invest considerable resources to get a person to even do so for part of the day.
How do you effectively steelman against your own position without a dynamic “adversary” responding to counter-claims? I feel like every static argument except the most absolutely resolute is vulnerable.
It is possible that Anthropic (and OpenAI for that matter) actually put out some pretty low quality software.
If you’ve used Claude Code for any length of time you’re familiar with all of the strange rendering bugs, freezes, etc. OpenAI is even worse, their horrific software makes it difficult to _pay_ them, which should be top priority for a company.
It's weird to me how so many people just put up with crappy, low quality software. If you buy a physical good and its defective, you return it, stop buying that brand, maybe even leave a negative review or contact consumer reports, etc. If it does not live up to what was advertised, you go to the FTC.
But when it comes to software, we all just kind of accept that shitty software is the norm and totally fine? Let's start calling it what it is, its defective.
What really grinds my gears is when the app is a _broken_ version of the website. I find the app really expects you to be on the happy path in a way the website does not.
Uhaul is a great example of this. The app looks a little better on mobile, but if you don’t make an account and search by reservation ID, all features related to things like extending a reservation time are nonfunctional. I have to go to the mobile site which looks bad and has poor navigation, trying to hijack my session and redirect me to the app, but at least it works.
UHaul is the only app I've ever encountered that is less efficient than talking to a human being behind a counter. I have tried multiple times to pick up a reservation via the app and it has taken so long or failed at some point that I had to walk into the store anyway.
And it wants a photo of my license EVERY GODDAMN TIME. And no I can't store a photo of my license on my phone and just upload it - it has to take a photo directly. And boy does it not like to take photos in bright sunlight. You know, when I'm outside in a parking lot beside a white truck.
That's funny, I've NEVER, not even once, encountered ANY app from any company, in any business, that is easier or more efficient than talking to a human being behind a counter. (So long as the human behind the counter is not also hamstrung by dreadful Indian crapware...)
Seriously, like most people, I have dozens of apps on my phone, forced on me by various needs. The ONLY one that makes anything easier or better is my password manager, which I have to have to safely keep and lookup the logins and passwords for all the damn apps...
I dunno. It's so hard to get a plain ordinary brewed coffee at Dunkin these days, the kind of coffee I used to get there back in the 1980s. On the app I could actually pick one out, in person I'm likely to have a teenager give me something else and then have a nervous breakdown when I point out the one thing I care about on their digital signs: "If we get your order wrong we'll re-do it"
I pretty much exclusively drink black coffee or plain espresso and I’ve yet to encounter any issues getting it.
I have seen lots of complaints and memes to this effect, but never has it been my lived experience. I’ve also never had an issue with politely saying “excuse me, I actually ordered X and I got Y by mistake.” No nervous breakdown at all
I grew up in New Hampshire where it is almost as much as a religion as it in Massachusetts. We had one in Ithaca for a short time when I was in grad school which didn't last very long, I remember sitting there while a schizophrenic woman told me that the mob ran absolutely everything in New York, it was closed a long time, then a Thai restaurant moved in and then, decades later, a Dunkin moved in next door.
Early on I learned that the P&C supermarket made better donuts so when a dunkin opened across the street I wasn't so interested. And of course Dunkin used to be known for good coffee but now you can find good coffee anyhwhere except for a Starbucks and Collegetown Bagels.
The car rental desk at airports is one of the worst customer experiences I've ever seen. It's sometimes worse than what seems possible. It's sometimes bad in ways that don't even make sense.
Their scheduling software is also blatantly broken. I understand the difficulty in monitoring a large distributed fleet, but their software loves to say trucks are unavailable where a quick phone call reveals they can accommodate the exact request.
> What really grinds my gears is when the app is a _broken_ version of the website. I find the app really expects you to be on the happy path in a way the website does not.
It reminds me of that oft touted experiment where a light lit up within some number of milliseconds of a person hitting a button, and people perceive it as happening before they hit the button.
Could you not say the same about using a compiler or higher level language or a library you don’t understand or an algorithm you don’t understand or a chip that you don’t understand?
Now think about what a typewriter does in that context.
Now think about what an LLM does in that context.
Can you reason about how those are vastly different tools within the context of writing or communication, so much so that the comparison doesn't actually make any sense?
Hint: it has nothing to do with determinism. It has to do with the nature of the work itself and the role of the person doing it.
As an aside, this whole exchange really is just a perfect encapsulation of the outcome-focused versus process-focused individual.
The outcome-focused person sees words in a document and whether an LLM produced them or a human typed them is a distinction without a difference.
The process-focused person is utterly baffled that anyone could think those are in any way equivalent.
HEH. I just read your comment three times and I definitely don't get it. I think its entire point is to express outrage about not getting it, rather than actually explain it.
Perhaps we'll get to a point where believing any un-sourced information from an LLM will feel crazy. I don't want my model to know more than it needs to perform logic and use tools. Once it is capable of using tools I would much rather it looked up information or sourced it from existing context rather than just divine it from it's weights.
The problem is it needs world knowledge to know what to lookup. This puts a floor on how little it can know while being able to look up what it doesn't know. Maybe its better if it knows a lot but has a good instinct for verifying that.
How do you actually get to the thing you're looking up? The scrapeable internet is shrinking in response to scrapers.
"Source: rare book ingested and shredded by Anthropic. No, you can't look it up and we can't show you the scan. The remaining open market copy is $5000. Trust me."
I empathize, and I have the same preference, but I wonder how this interacts with other people (many of them being our coworkers) using LLMs. There is no authoritative source for the models to pull info from, so either people will have to exercise good judgement and double check important claims, or they will trust too blindly and fall close to the level of whatever LLM they use. In that case, I prefer my coworkers to use an LLM that does have world-knowledge -- I will still hear them spout ridiculous claims, but at least it should be less frequent. It strikes me there's a sort of prisoners dilemma here, where if nobody trusts others to critically evaluate info, it's in our interest to make the tooling do it instead, to whatever degree that is possible. Maybe I'm too cynical about working with others though.
Probably. You can solve it with either some grounding context, or spending hundreds or thousands a month extra on a model that has more knowledge baked in. With modern harnesses, the choices is obvious.
Doesn't matter if you aren't asking the type of questions where hallucinations are relevant e.g. you're seeking pure reasoning rather than factual information.
reply