Hacker Newsnew | past | comments | ask | show | jobs | submit | akersten's commentslogin

It's useful because it let's me see the decisions the model will make before it wastes a ton of time implementing them. The model is smarter now but that doesn't solve for underspecification if it guesses my intent wrong

Interesting, I don’t see this very often with the latest models. Are you using Opus 5.5/Fable 5.1?

Either way, plan mode isn’t going away. You can always /plan or ask Claude to enter plan mode. We might re-map the shift+tab keyboard shortcut to something else by default for people that don’t use plan mode.


> Interesting, I don’t see this very often with the latest models.

Strange response. I agree with the parent comment here, plan mode lets me ensure that I have specified everything correctly before it gets built which is far too late. I don't see how an improved models even matter to this workflow. Is Fable going to read my mind?


What is it with the snide responses? This threads title is "Plan mode is dead", not "You don't need to plan anymore".

Boris is saying that you don't need /plan to get the model to plan, you can just say "Let's plan this out" or similar, which at least matches my experience. Your experience may differ, of course, but it's not even clear we are talking about the same thing.

Maybe some people have not been long enough on this rodeo: This used to be an actual issue. You told the model "DONT START CODING YET" and yet, surely enough, starting to code it did. That is what /plan etc were supposed to fix.


I also use Plan Mode. How does it otherwise work if I am not even quite sure what exactly I want to build myself?

> if I am not even quite sure what exactly I want to build myself

Sounded to me like you need a plan.

My approach is to take the statement of work or problem definition and iterate on that myself until I'm really clear on what is the goal. I therefore have a good some good ideas about what the plan should be.

If extending an existing application, which is usually the case, then make use of the plan documents that I had written before AI arrived on the scene. These sre documents in markdown form that say step-by-step how to, for example, add a new report to the system.


You tell it, "let's brainstorm, don't implement anything yet". Then you talk about what you want to build and hammer out all the corner cases. Then you tell it "now do that".

I have something like this in my default rules which Claude has consistently always loaded before doing any work. Works really well!

> How does it otherwise work if I am not even quite sure what exactly I want to build myself?

It makes some plausible choices and you can retroactively ask it to make different ones later, if you want.


My experience with that is not very good. It gets so hung up in its initial decision. Like, it won't make changes because they're "breaking", for something never even committed. Or it will litter the code base with defensive code and comments about the path not chosen.

This is my experience as well. Changing the initial decision is a hassle so I just start over most of the time.

> My experience with that is not very good. It gets so hung up in its initial decision. Like, it won't make changes because they're "breaking", for something never even committed.

This hasn't been my experience.

> Or it will litter the code base with defensive code and comments about the path not chosen.

I've definitely seen this, though.


Edit: Deleted.

I love when people on HN think they're special or do more "real" software engineering.

Basically every big tech company maintains "codebases with millions of lines built upon decades." Talk to your friends at a FAANG and ask them how they're using Claude/Codex.

Merging after reading the PR description is just how it's done these days, and if you can't do it reliably, your harness, devloop, or model is simply behind the times.


>when people on HN think they're special or do more "real" software engineering.

But you're doing the same with your "how it's done these days".

These days things are done in many contradicting ways, and it probably will take at least a few years to settle on common normal.


The model is irrelevant, plan mode helps us ensure the model will actually do what we have in mind. The best model in the world can't work around misinterpretation because of bad specs. I'm not sure why this is even a discussion, isn't it obvious to anyone actually using LLMs?

This is an out of touch with reality thing to say. Remember: not everyone has access to unlimited usage/money to spend tokens or the most expensive models with the higher thinking settings, exceptions does not make the rules.

> This is an out of touch with reality thing to say

I mean, you're discussing this with a marketer/someone wearing a marketing hat, who works for a company which needs people to use as many tokens a possible. That's their reality


[flagged]


"Marketing hat" is extremely civil. Back off.

I see you didn't disagree with the main thrust of my post.

You'll also note their 'About' is empty. Regardless, engaging with HN in this manner is de facto marketing/PR. Trillion dollar companies don't just let anyone post on high-profile social media websites for any length of time without permission from marketing/legal/PR.


Engineers can nevertheless also be marketers for their employer's products.

Hey, well first off, congrats on making the greatest product ever probably.

I absolutely see fable and opus 5.5 misunderstanding intent, but that just seems to be a feature of necessarily underspecifying in a written prompt. Just today, I gave opus 5.5 a simple task to spin up a new environment for work. It read the ticket, which was decently specified and knowing the codebase as well as "Ghasp... reading the code" I had to correct it about 5 times to do it in a way that I would have expected it to. Getting the pipelines right, environment variables, and configs. It was all relatively straight forward imo. Then I had to prompt it to clean up its corrections, because it left a workflow variable in the github action that some intermediate step required but the final solution didn't. I definitely would not have caught that if I didn't read the output. Idk, there seems to be a natural limit as to how much it can infer and I have no idea how to fix it. I did write about it [here](https://javiergonzalez.io/blog/the-assumption-problem/) though.


I’ve been having similar issues. Absolute love fable but it keeps leaving development servers running that are blocking port 3000 (rails apps run on this by default) and then when I try to launch the app and realize the port is in use I ask fable what’s up and it says sorry I left x running and then shuts it off freeing up the port.

Theo would have a field day with this, should surely read your article too imo.

Why would he have a field day with it? It seems he oscillates between finding models the greatest thing there is to being absolutely stupid so not sure in which direction you're implying he'd have a field day.

I'm curious. Who is Theo?

Theo is a YouTuber & prolific X tweeter who has a channel about AI. This is not Theo, but gives a flavour of Theo's style: https://youtu.be/h1p9zdUtUdo

thanks for the info. I'll politely pass on the video though. priorities

I don't think me giving the model bad instructions is something a smarter model can solve. I use plan mode constantly (with opus/fable), and at least once a day I'll say something too vague or just dumb and it will sketch out the "wrong" solution in it's plan.

Which is fine because it just put together a plan and didn't spend 10 minutes rearchitecting everything.


I use fable 5.1 (tried opus but it lied to me 3 times in quick succession and ignored me in another)

This is weird to ask because I feel like of course the model isn't omniscient? Isn't the whole point of iterating on a plan to assess impact, risk, know your (the user) variables, user impact, product impact, etc for making a change? I cannot count the times even in the past few months where I start a conversation with my C suite because their desired outcome would have a potential negative impact elsewhere for other products or users.

Is this just not something that comes up at Anthropic?


It happens with Opus when I leave too much up to interpretation and the AI doesn't do what I had envisioned but didn't specify. Like sure, what it did may be a technically correct solution but it's not the correct solution that allows for further development of my idea. I'm not sure how others do their projects but I start small with proof of concepts and develop in layers until the project does what I want. I use plan mode first to layout everything I can think of that I ultimately want and describe features in the best detail I can manage. I work with Claude to figure out the best framework or find whatever existing projects can serve as a starting point. The first milestone is the proof of concept, take the framework/existing project and build something that does the bare minimum of what I need in the way I want it done then build a test suite to make sure it works. Once that's proven out, we start adding more features (both mine and the ones Claude has suggested) and adding/revising tests along the way. For small things, I won't bother with a plan since I generally already know what I want or any ambiguities can be solved in a single response. But for larger things, I try to take a waterfall approach with well defined milestones.

If I knew exactly how I was going to build something, I would have built it myself. But since there's some ambiguity in the portions of the project I'm less familiar with, I rely on the plan to not only help me understand the decisions Claude has made for me but to keep Claude constrained to the decisions I've made. It's very frustrating to waste tokens on having to refactor something because


I can relate to this. But this part doesn't make sense to me.

> But for larger things, I try to take a waterfall approach with well defined milestones.

> If I knew exactly how I was going to build something, I would have built it myself

Aren't these contradictory? If you don't know exactly what/how to build, how can you do waterfall?


Yeah IMO there's two entirely different traps I've seen in startup codebases in the last year or so:

1) "Maybe waterfall works now" - plan mode, take care of all the nits and issues that the bot leaves on your PR, wildly overengineered "enterprisey" solutions with a lot of bells and whistles all over the place but very poor end-to-end user story test coverage that results in user experiences with a lot of good test coverage of the edge cases of how a given step might fail but little thought towards overall user flow and throughput. Because part of the issue with waterfall was assuming you could design the right tool for your user up front.

2) "Maybe code doesn't matter anymore" - The just ask for something when you need it approach, which results in weird janky individually-sorta-working but strangely-overlapping six-variants-of-the-same-thing that makes it hard for your users to develop a single consistent mental model of the thing they're using, and that changes super frequently.

They both end up with a lot of other bad-for-velocity things that I assume are inherent to how the tools have been refined in response to last year's criticism, too. Super verbose comments. Extensive - without much eye toward runtime - low-level test coverage that might miss the forest for the trees and also slows down the next round of iteration cycles. A plethora of new proper nouns all over the place that make the documentation an ouroboros without a good entry point.


I haven’t tested with Claude specifically in a while, but I see this a lot on larger features.

It tends to be small decisions way down the stack that bubble up, or an incoherent data model that can’t handle what you’re asking for cleanly.

Eg I was messing with a state tracker the other day. The state tracker assumes a container is either currently running, or fully removed from disk.

The LLM chose to remove the state file when the container is stopped and then to remove it after, which leaks container storage.

The LLM is kind of stuck though, because every option other than “rewrite the data model” has negative outcomes and it probably violates user expectations to launch a massive rewrite there.


I've found Opus 5.5 is quite good at surfacing and clarifying these issues using Grilling[0]. Often I find that I want to expand the scope of work much larger than Claude would consider based on my original request.

0. https://github.com/mattpocock/skills/blob/main/skills/produc...


We got Opus 5.5 just a few days ago so haven’t had time to learn its ways, and Fable was/is too expensive for many people including myself. Insiders have months to learn the ins and outs of new models.

It would be nice if with new model releases claude code also gave a bit of a model 101 that tells you evolving ways of prompting it that the insiders have picked up. I know there’s the prompting guide in the claude docs, but this is often very broad and most people don’t know about it.


I have primarily used Fable 5, Fable 5.1, and now Opus 5.5. I never use shift-tab to enter plan mode personally. My general workflow is chatting with the machine to work out fiddly bits, then /plan to get everything in one document, and reviewing the plan.

/plan is still useful, I still need to review what it's going to do and still make revisions. But there's two phases: hammer out key design decisions then write and amend the document.


For me, every day, although it fails far less often at this than Opus 5 did, which might as well have been a gremlin. It was a really poor technical writer too, leaving slop reminders to itself in all prose, including comments. So even if 5.5 fails less right this second, I have trouble believing that it all will not similarly bite me next week.

Summaries that don't tell me when it's changed direction in a timely fashion, but I am only told way later, when I have to undo. Really bad judgement calls regarding where to fix bugs, changes in implementation decisions, taking action when I am asking a question directly, not passive aggressively asking for action... when 5, 3 days ago, was proven to be untrustworthy, switching to very little supervision sounds like a strange thing for a customer to do.


Why is “are you using the latest model?” the automatic response to any even mildly critical of LLM coding?

OP wasn’t even actually critical of LLMs, they were just saying that plan mode was helpful to stop the model from making incorrect assumptions when you want you don’t specify everything you should.


How could a newer model be better at making up information? Do you understand basic information theory? Or maybe what you're saying is you don't actually have any ideas and are happy to do what everyone else has done before you so it makes no difference?

It’s just surprising how few people understand this. It’s not like it’s new either. Polyani’s theory of tacit knowledge captured this back in 1958

The older models were less good at inferring intent. The new ones just seem to do a better job.

They're just more subtly wrong when they are

That sounds like a contorted way of saying "better."

Hi! Taking this moment to gripe; forgive me...

5.5 and 5.1 have major Rain Man (savant) syndrome. Excellent at many hyper-technical things, absofuckinglutely boneheaded at anything that a human (or an earlier model) would understand - like how to write copy, what a human would expect in a given situation, various types of norms...

it's infuriating because it's a sophies choice - dumber model but better human understanding, or better technical model that you have to explain things to over and over like a toddler.


You can just tell it to write out a plan.md file.

I greatly prefer this, since it lets me iterate on the plan with Claude for a while without it repeatedly asking if I’m ready to implement the plan.

Once I’m satisfied, I usually start a fresh session and tell it to implement the plan.

For smaller plans, you don’t need the file. Just ask it to come up with a plan. I don’t recall the last time it just started implementing if I only asked for a plan.


This is the process I generally use too. Small plans you can just ask for, and big plans you work through building a plan.md file before you build it.

That also makes it easier to adversarially review the plan (I have Fable write the plan, then review it with Astra and another Fable instance).

But that’s just the value of planning, not having it be a special mode.

I guess that You dont need plan mode any more is ambiguous. I took it to mean you dont need to plan because the models are so good at infering intent. You are taking it as "you dont need a dedicated modality to create a plan".

> wastes a ton of time

It wastes a ton of tokens as well and those are not cheap.


It’s not that planning is dead, but rather that planning has outgrown the simple “Plan Mode” feature as models have become capable of taking longer turns.

the thing is it fails on CSS most of time, I experienced it and it takes a lot of time to fix again and again, and ruin the code sometimes

If your core service is getting more expensive to provide and competitors are busy eating your margins, why let someone else taste your secret sauce and only get paid for the tokens, when you can keep the good stuff (bio capability) for yourself, and net both the profit and the fame?

> When the unit doesn’t work correctly, the anger and RMA requests are directed back at the Raspberry Pi foundation.

They don't need the resellers help for that! I've already been burned twice thinking "I'll use a raspberry pi for this small project" and having it die within a month because reliably reading/writing to an SD card is hard or something.

Between the notorious unreliability, crazy price hikes, trademark or whatever disputes, and this anti consumer ewaste generating policy, I can't think of any reason I'd ever patronize them again.. my only regret is it took me two purchases to learn this lesson!


Their sample prompt takes me back a year or two:

> Find unused code in our web app. Exclude generated files and test fixtures, and check for indirect references before recommending a removal.

If I "/goal remove unused code" in Claude today, I would not even think to specify "check for indirect references" and "don't consider generated code dead code," those sorts of intuitions have been "built in" to the frontier models for a while now.

So I'm pretty confused by this, what does it do that I can't with an agent swarm?


> I would have guessed that reliably identifying LLM generated text was not possible

It depends what you mean by "reliably." If you mean, "we should be comfortable relying on this kind of tool at scale to identify and punish students, professionals, and writers who may have used AI," absolutely not.

If you take "reliably" to mean "1 in 200 false positive rate" as they disclose on their front page, absolutely that is possible (they are doing it today!). If you think there are more than 200 assignments turned in over a given year at university, you probably do not consider a tool like this fit for purpose. It's an open question whether those procuring said tool are aware of this

Unfortunately their marketing is really insisting on the former, and trying to push it into the zeitgeist that detection of AI-generated or edited text is reliable-type-1 now and long-term. They fail to make it clear that this is merely a tool that strongly suggests text follows patterns known to us at the present time of known LLMs. However, that fingerprint will drift over time, as LLMs get better, human writing style evolves, and the line between human and "smart autocorrect" becomes even blurrier (does speech-to-text push the model into "AI assisted" mode, because it tidied up your punctuation, for example?)

"What color are your bits" is good reading today as it was 20 years ago: https://ansuz.sooke.bc.ca/entry/23


> If you take "reliably" to mean "1 in 200 false positive rate" as they disclose on their front page

Pangram claims a 1 in 10,000 false positive rate (rate at which human-authored texts are incorrectly classified as AI-generated). 1 in 200 sounds like the false negative rate (rate at which AI-generated texts are classified as human-authored), or perhaps a rate for a specific category of text.


Good point of precision, I read their site too quickly. I don't think my argument materially changes with that number instead however. There were over 40,000 students enrolled in my university alone. Generously assuming they only turn in one assignment per year, having 4 of them go through the "computer says you cheated, and as you can see, it's 99.9% accurate" gauntlet is not a price we should be willing to pay for... the marginal benefit of this tool over more classical ways to proctor and assess pupils

> maybe it’s better than nothing?

While that is not quite my bar of confidence when implementing wide-reaching technologies that have numerous unexplored knock-on effects, I guess the calculus must have been different on Infinite Loop recently.


The fundamental issue isn't technical. It's that people will see the "certified real" tag and just take the image for face value of whatever narrative someone wants to convey. They'll see the "Real Photo, Verified by Apple" and their brain will short circuit [0]

I don't think we should have this, for that reason alone (but many others too).

[0]: https://imgur.com/fVPkpuQ


I’m pretty sure “certified real” aren’t the words Apple will use, nor do they use it in this document. The words to describe the technology were chosen with care: semantic verification, attestation, tamper evident, etc.

But that’s not how the label will be interpreted in real life

You seem to be missing the point that’s being made here. Of course apple will word this very carefully. This does not matter in how people will interpret it. They will interpret it as certified real.

The average person already does this with obvious AI slop.

Why are you subjecting yourself to that instead of downloading Firefox + uBlock origin?

See the headline of this article for Exhibit (I've actually run out of letters in the alphabet) why adblocking is a moral imperative.


As much I love a good old "Why don't you just ..." response that terrifically misses the point - in this particular case I was watching YouTube through their Apple TV app.

Now, there may be another "why don't you just ..." or "well, actually ..." response you have queued up, maybe some ramblings about a pi hole, or some anti-Apple hate, or some other solutioneering. Go ahead, tell me.

But it is the structure of incentive that forces this. The YouTube creators that make the content do need to get paid. And Google offers an option: Youtube Premium. That give both me and the creator what we want and (for the time being) removes the ads.

So either one is rich in time (doing whatever "why don't you just ..." hoop you expect me to jump through) or rich in money.


If you want creators to get paid, pay them (channel memberships, patreon, click on their sponsor links and buy something). No need to subject yourself to ads.

Ads don't just steal your time either, they invade your head. You are influenced by them and you have no choice as long as you watch them.


Sounds like you've made a choice to watch ads. Which is fine, but then maybe don't complain about it.


That is wrong on both accounts. I have little choice other than to watch ads since the structure of incentive that creates a marketplace like YouTube forces it on me.

I think about libraries and how the world would be if instead of a free public resource, you had to watch ads before you were allowed to take out a book.

It is just the case that in this modern social media world, if you want your ideas to spread you have to put them onto social media. And if I want access to those ideas I have to consume them through social media. There is nothing about that environment that is my choice, it is the environment that I find myself in.

And I have few choices to get around it. I can twist up into a pretzel trying to block the onslaught with more "why don't you just..." advice. Or I could pay to make it go away. Or I could just go off-grid and forego access.

And when you look close at the options, it becomes clear that sovereignty of my own mind has a price. Freedom of mind is not free in this modern world, if it ever really was.


You made the choice to use Apple TV, whatever that is. I just use... YouTube. In my browser. Do you not have a browser? It doesn't cost money.

You also have no obligation to watch ads for creators to get paid. They can also set up things like Patreon if they would like. The ads are very lucrative, and that's their choice. But you are free to make choices too, if you would like to.


They wanted a whinge, and not a solution. A position that I truly do not understand, but is not uncommon.


> I was watching YouTube through their Apple TV app.

That is an active choice, not forced in any way.

> The YouTube creators that make the content do need to get paid.

There are better ways, and we would all be better off except for the small group of greedheads benefiting from the chaos

The ad supported internet is dying. Good riddance


>The YouTube creators that make the content do need to get paid.

The world would honestly be better if they didnt.


> The YouTube creators that make the content do need to get paid

Nope.

They may get paid, they may not. That's the risk they've chosen. They may also get their Google account banned with no recourse for something minor or even non-existent.

> So either one is rich in time (doing whatever "why don't you just ..." hoop you expect me to jump through) or rich in money.

That's what they want you to think, you've been watching too many ads.


We must ensure the gravy train keeps rolling until we IPO.

> Crack down on unauthorized distillation / prevent weight theft

Actually hilarious to put that in writing, given the genesis of this entire business model.


Pulling up the klepto-ladder.


Eh, I hope Apple continues to provide this and it forces the needed discussion about how two-party consent requirements are nonsensical. Why should it be illegal for me to remember exactly a conversation that I participated in, instead of only being allowed to have a vague recollection?

Laws like this provide cover for abusers and deceivers, by preemptively spoiling objective evidence and making any accusations depend on hearsay instead.


It's not illegal for you to remember, it's illegal for you to record. There's a difference.

For me, I don't want to live in a world where if I say something embarrassing or not well thought out, someone pushes a button and the last 15 seconds is transcribed as evidence. That's a different world to the previous one where it would be someone's word fori it.

As for the fact that phone could already do this, that's not the point. Phone users have to go out of their way to make it happen so it's generally unlikely to happen. The watch feature though is always on and just waiting for you to press "save last 15 seconds"


Same with the always-on transcription feature that you dont even need to interact with. I like not having to speak with extreme precision, knowing my words are going into some record.


> Eh, I hope Apple continues to provide this and it forces the needed discussion about how two-party consent requirements are nonsensical

Really disagree on this, I think all states should be two party consent, personally.

> Why should it be illegal for me to remember exactly a conversation that I participated in, instead of only being allowed to have a vague recollection?

It's not illegal, it's just that the other person has to know that you're doing that and consent to it.

Giving them the chance to walk away or to tell a person and their Meta glasses to fuck off is important.


People are fucking liars. Everyone, everywhere, always.

I want to record every conversation surreptitiously, because that’s the only way to catch them.

The reason we can’t is because our politicians and legislators are the biggest liars of them all. It would spell their immediate downfall.


Our politicians and legislators lie on camera constantly. But it doesn't really seem to matter much these days. So I'm not sure why this would make a difference.


Everyone being liars means you and I are both liars, same as the politicians. But the politicians have the power and if we remove two party consent then they get to surreptitiously record you and leverage that power. It doesnt even the power playing field.

We can alreay write notes down for every conversation and then send them to the person involved saying "we talked about X, Y, and Z." That last step is the key because it lets them object in writing if you mischaracterize things. From a "catching someone in a lie" the most important step is that one, because you form a paper trail where the other party can correct or contest what was written and bring that up now, and the fact that they didn't is itself evidence in case of a dispute later. The apple watch feature doesn't do that, it just dragnets everything. Even if it recorded the audio, we are in a faked-audio world so unless you have some signal that they agreed that they said a thing ahead of time, they can always deny it later.


Yes, let’s do something illegal to force a discussion. A discussion that already happened when writing the law…


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: