Hacker Newsnew | past | comments | ask | show | jobs | submit | hedgehog's commentslogin

I think you underestimate how intensive it is to scroll LinkedIn in Edge while a YouTube video, Outlook, Adobe's updaters, antivirus, device management, and whatever incidental malware the user has picked up run in the background. That's the reality for a lot of people.

I believe that falls under competent software. However, I do see your point.

How did you organize the atlas? I've tried a few things including embeddings and clustering files based on how often they change in the same commits, haven't yet found anything I want to bake into my tooling.

I've gotten good usage out of a structured search tool (not mine, by someone else here) called Tilth:

https://github.com/jahala/tilth

It combines search with tree sitter grammars so the results can annotate usage vs definitions, cite line number ranges, inline the actual definition if it's short, etc. Not as precise as LSP but simple (no daemon), human readable, and in many cases works without configuration.


It's useful for search type problems where the result is relatively compact and verifiable in isolation. I haven't found a way to build production quality software that way in general but for certain kinds of reverse engineering or system optimization problems it's helpful.

Are you talking problems along the lines of Karpathy's autoresearch? Reverse engineering does seem pretty straightforward actually in that use case, but for practical business building work as opposed to hobby projects that's where I'm hitting the wall at the moment.

I haven't looked closely enough at his autoresearch stuff to know how similar it is. From what I have seen the quality of output is directly tied to the quality of the request, a human can only read and understand some finite amount of material in a day, so maintaining enough understanding to ask quality questions becomes the limiting factor. What are the priorities you want to move faster on?

There is some research suggesting that a prefix from a stronger model will tend to elicit better completions from a smaller one. I am doing some experiments to see if I can replicate this in a practically useful way, e.g. Fable + 4B Qwen, or 125B Qwen Flash Next + 4B Qwen, results TBD.

They confused "smaller" with "budget" though, if the good camera was available in the smaller size I would have bought it.

Camera blocks take internal volume. Smaller phones have less internal volume to spend - while still having to pack all the non-negotiables like the modem and the SoC into it. Something's got to give. And no one want that something to be the battery life going down to 6 hours.

Volume constraints are bad enough for "normal" models. "Minis" have it way worse.


The iPhone Air's SOC demonstrates that they could compact the actual SOC portion substantially. Add a bit of thickness back for the battery, and they could have a Mini with space for a decent camera and decent battery life. Remove the camera bump by making it a uniform thickness and the battery life would be fine.

But we already know that they can make a smaller iPhone with everything.

The iPhone 6 was a decade ago and was smaller in all dimensions than an iPhone 17 and had a perfectly acceptable battery life. Since then, they have removed headphone jacks and buttons, batteries have gotten more power dense, chips have gotten more efficient, etc.


They used to be able to. I don't think they can anymore. The iPhone 6 was 3 years after Jobs died. Apple's corporate/engineering culture has had 12 additional years to rot since then. If the Scott Forstall situation, Chris Lattner situation, etc, tells you anything, it's that Apple cannot make basic, common-sense tradeoffs anymore. All they can do is copy whatever the wider industry is doing.

All of the excuses people are coming up with in these comments are just that: excuses. Look at the sibling comment by musictubes, for example. Citing these ethereal "critics" and "most people" when in reality Walt Mossberg was the only critic whose opinion mattered.

Occam's Razor: The Innovator's Dilemma. It's a rotting, bloated megacorp with too much to lose, that doesn't take risks, and can't change direction. The direction it's in is "bigger phones", for whatever reason. So we're gonna keep getting bigger phones. It's that simple.


It’s the camera and Face ID. The camera on the iPhone 6 is nowhere near as capable current cameras. Even the much improved camera on the last iPhone SE was considered a huge issue by most people. Cameras are too important to go back to something that thin. The current Air is leaps and bounds better in every way to the iPhone 6 but it doesn’t sell. Why? The battery life and camera. Imagine shrinking the iPhone Air and having even less battery life. Critics would howl and even fewer people would buy it.

Smaller phones can’t fit as much stuff in them as larger phones. Shrink the phone and something has to give. People in the aggregate do not want to make those compromises unless the price is noticeably less. But then it can’t have any of the premium features that are expected from Apple. Small phones are compromises that don’t pay off for Apple.


We’ve come close to doubling battery capacity per volume in that time.

The battery was the single largest module by a long shot in an iPhone 6, and now you can get the same capacity in a much smaller volume. We’ve also lost other bulky items like the home button and headphone jack.

The camera on the current gen has its own expanded volume area of the case, so I don’t buy that they couldn’t just use the same hump they use on literally all of their phones. Sure Face ID is new, but I’ve seen the size of that module, and it is only slightly larger than the selfie camera from back then that it replaced.


Reading this on my 6s, and it’s still great.

Same, I look longingly at the Samsung Flip and I'd buy one in a heartbeat if only it had the camera system from the Ultra. I'll even pay the Ultra price+extra for the foldable function.

But right now it's in a space where you're paying extra for the fact that it's a foldable, but it feels....budget in its specs? Like...how does that work?


> They confused "smaller" with "budget" though

Apple is in the luxury business.


I miss the time when smaller was considered more advanced

There's something funky with the renderer, it looks like perspective is wrong. All vibe coded?

Can you link your scene? It's possible there is an edge case.

The one you linked here in the comments.

I think you might be talking about how the camera uses orthographic rendering? This article explains it well https://manual.keyshot.com/keyshot11/manual/cameras/lens-set...

That's a good reference for what orthographic rendering is, but if you look at the rendered scene it's definitely not orthographic. Parallel lines in the model converge towards the camera when the model is viewed from an angle.

I think it's an optical illusion. It looks weird to me, but comparing it with various objects I have I do feel like they're parallel.

Specifically the roof edge vs base edge, the back of the roof vs the front of the roof.


Probably better to use perspective as the default like Sketchup.

This roughly lines up with my personal experience that in March a combination of stronger models and better tooling on my end let me start running jobs unattended 24/7 (using Anthropic sub and my own hardware). Their $8000/day per researcher spend is crazy though, I'm curious how they keep track of the work.

Sounds like OpenAI are in the token-maxxing camp, so who knows what individual employees are doing to work their way up the leaderboard?

If you spend $8000 to generate an animated pelican riding a bike, then how much tracking does it really need?

Is the guy who spent $300,000 or so translating the FLT proof to Lean going to get a big Christmas bonus?


That was Anthropic.

End of day, output and results are top target of measurements, token consumption is the obvious number that they would like to disclose for their own business benefits and a simple metrics that correlate with the output.

Rest assured, capitalist appears irrational in wasting money, but they certainly care more about profit.


Taking a profit means you have to show numbers and the sooner you show numbers the harder it is to take people’s money.

Can you elaborate on this? Especially the tooling.

I tried something similar and I remember it was still pretty dodgy in February.


my stack in a sentence: refine the docs/prompts/skills often, that's your biggest job, use both frontier labs models reviewing each other, don't solve individual problems only the systemic ones (set standards strategically, don't define tactics)

If I had that many tokens/dollars I would be running canaries and adversarial verification in prod based on e.g. traffic replay, live fuzzing, all kinds of things to build confidence without direct human line-by-line review. If I had $100k to spend next month I could probably get through it, I'm running $2500+-api-equivalent a week at this point and I feel very token limited. Will be time for a 2nd or 3rd subscription soon for both labs I think.

Fable was a revolution, still learning how best to use it, 5.1 felt like a notable upgrade. At this point I launch a workflow with 10-20 minutes of interactive setup (and even that I feel might be too much), it runs for hours, and the PR is trivially mergeable (I still review every line, but 95% are just merge, maybe 4% are feedback needed, 1% are thrown away and regenerated, which implies I'm being insufficiently ambitious)


Thanks!

> If I had that many tokens/dollars I would be running canaries and adversarial verification in prod based on e.g. traffic replay, live fuzzing, all kinds of things to build confidence without direct human line-by-line review.

This part jumped out at me. There's something to watch out for here.

I recently had a funny experience. I delegated a major feature to an agent.

It turned out that it had implemented it precisely backwards, in a way which was pointless and which made things worse.

But it had written countless tests for the feature and all the tests were green.

I realised in that moment that even formal verification would not have helped, because it would simply have written a mathematical proof of the correctness of the incorrect feature...


Yes, you need some kind of other source of truth. I think the best way to get that is to do clean room development with a different agent, but ultimately if you give them the wrong idea they'll do the wrong thing.

The other thing I do, not as much as I should, but it's very powerful, is to generate spikes and deliberately throw them away to understand how to prompt better. Like I generated a swift version of the react native app I'm working on, and Alloy provers for the state transitions. None of it is production quality but getting great results that way is useful to scope future work.


If you've managed people, these are all familiar problems. I found you need much more than a functional specification, you also need motivation, background, related work, ideas tried, etc., because those help disambiguate the right path in the inevitable situation where your original task description is unclear or conflicts with itself.

Yes I'm using the entire consultancy stack - define values, etc, and work your way down the "where do these not match reality on the ground and need change", but for little robot people instead of (arguably less messy) humans.

> I think the best way to get that is to do clean room development with a different agent

What is this about? Could you give an example?


So, if I'm building, say an iOS app, I have a) the code itself (swift/react native), b) the internal test suite (unit tests), and c) the e2e test suite with something like XCUITest that drives the physical development phone, plus d) the backend (which has similar tests but is basically a mirror, so I won't elaborate)

So when you're creating a (and hopefully b), you use one model with one context, and then you might use some other model to implement c. I kind of round robin the models and present them with seperate context - so for example I don't build a, b, and c together, I build a plus some b, then later go for a pass over b and c together, then maybe I used c to drive improvements in a, and I vary between openai and anthropic models as I do.

Heres an example set of minimal prompts:

a-focus:

implement feature <x> based on #ticket in github, be sure to reference the engineering standards documents and the swift and react native skills as needed

b-focus:

improve test coverage in the repo for <subsystem z> to ensure that <feature x> is covered completely, and fix any outstanding gaps or omissions in that feature as you go (standard references above)

c-focus (possibly in a separate repo):

You are creating a black box xcuitest to drive a physical phone for testing <feature x>, here is the user specification and known issues, create failing tests for each known issue and an overall robust suite to ensure any user facing or ui issues are caught


Pretty much goal + task + dependency infrastructure to help avoid drift during long runs, especially across compaction boundaries. I have spent a lot of time doing automation with models at various strengths including some of the early open-weights models (Llama, Mistral, etc) so I have a pretty good feel for how to steer productively, there isn't any deep magic just scaffolding built out of reading a lot of traces and debugging stuck agents.

That would be the mother of all circular accounting: the main clients of OpenAI are OpenAI employees.

These researchers are paid millions of dollars for their work. I doubt trust is really an issue at that level.

Yes, because no employee with million-dollar comp has ever been untrustworthy in the history of business.

Imagine if one of the humans at OpenAI was misaligned! We should get the AI to research this possibility once they've been aligned.

> let me start running jobs unattended 24/7 (using Anthropic sub and my own hardware)

How are you running jobs unattended 24/7 without hitting your token limits?


Similar to hgoel the model is managing the project's arc and writing + debugging code, but the underlying work is pretty compute intensive and all LLM output that is part of the final product is generated by local LLMs. Claude Code builds the pipeline that does the work, the pipeline runs fully on open-weights models and is reproducible top to bottom. There is a lot of detail to cost management. First is of course the Claude Code subscription is heavily discounted vs API costs. Then managing context size and turn count, which multiplied are basically what determine usage accounting (cached read is almost all of the cost). Auto compaction at 175k or 200k tokens (model the right number for your work), sub-agents with good model selection, tools to predict subtask difficulty so the sub-agents are correctly sized to complete under the compaction limit. Lots of focus on tooling to improve turn efficiency (e.g. the tilth utility by another user here for querying code). This started as a few scripts in one of my research projects but now is how I run all of my agent coding workspaces, and in another month will probably start replacing Claude Code itself for my purposes.

It depends on the time the job itself takes. If you're having the LLM handle a training run for another model, the LLM is probably spending most of its time waiting for iterations rather than consuming tokens.

For a task I left a local model running on overnight, only ~100k tokens were used because most of the time was just waiting on tests to finish, then waking up, tweaking a few settings and trying again.


Hi!! This is basically why I built Observer (https://github.com/Roy3838/Observer, FOSS). It runs a small local model that watches your screen/terminal and pings or calls you when the thing you're waiting for actually happens: tests finished, run crashed, agent stuck on a prompt.

So an overnight loop like yours wakes you only when it needs a decision, instead of you waking on a timer to check.

But in this case it would be triple LLM inception, one training another and a third one monitoring everything is done correctly :p


My only experience in >24h agents is with economically sane models (one of GLM5.2, 5.3-flash for orchestration, DSV4-flash for implementation, and glm5.3|sol|kimi3 agents + subagents reviewing at the end)

Over 24h my token spend is <30$. Excluding tokens for review it's <10$. With the absurdly gigantic subscription subsidies and a reasonable workflow I suspect one could run parallel agents.

I'm not sure what the point would be though unless working on some kind of optimization problem -- it takes me days to review <24h of the agent's output. It's almost always near enough to correct to be shippable; though I do give it feedback and iterate until it's better than the code I would have written.


It only really makes sense for problems that are complex and require iterations that don't themselves require much review. E.g. if you want find, PoC, and patch bugs, the output can be reviewed without reading all the traces. Or if you want to write a custom tool that does some job using local LLMs, assembling that pipeline, tuning the prompts, etc takes a long time but reading the final tests + eval data + code is enough to get a lot of confidence that it works right. Model checkers can help too, for example I wanted multi-sink Bluetooth audio support in Gnome for my kids so I hooked the hardware up and robo-coded the core logic specifically to be checkable with Kani.

>don't themselves require much review

I'm still wary of any unreviewed code - though my area of work is not tolerant of defects.

Agree on targets / verifiable indications of progress or success being a prerequisite for this being useful - although that covers quite a lot of SWE work.


This sounds like more work than just writing the code yourself. You'll say it isn't. I don't believe you.

That is an incorrect presumption - I think it's plausible this is more work; it's certainly far more taxing.

I'm at a point in my career where a small minority of my time is coding. The AIs can do in a day what would have taken me a week uninterrupted with acceptable (in some cases inferior prior to human feedback--but in some cases superior!) quality.

As I do not have 10 let alone 40 hours per week to devote to coding I think it increases the amount of high quality work product I can create with a given time investment. As I review it I merge small independent units and decompose the work.

All that is to say I don't really like it - but I suspect for most *well defined* coding tasks human produced code from highly experienced engineers will largely cease to exist in the next year -- getting cheap/relatively horrible models to produce good code is now straightforward.

OTOH I never use AI for any human facing communication outside of making my writing shorter. IMO AI slop "documents" are almost certainly a drag on organizational productivity.


I'm currently running two 24/7 semi-autonomous AI research projects using Fable 5.1. It's on track to burn through my weekly quota in about 3 days. I check progress in the morning and in the evening, and provide some light steering.

See my sibling comment, you can probably robo-code yourself some tooling to alleviate a lot of that in a few hours but if you want help shoot me an e-mail. I'm interested in seeing other people's workflows.

/loop ?

I suspect the $8000/day figure is the equivalent in API costs. But I also suspect gross margin on their API rates are 80-90%

Check sampling parameters and chat template, make sure you have adequate context window, turn reasoning effort down. It should be able to one shot a small app without intervention.

This is fantastic. It would be interesting to have Claude Code export an engineering guide for doing similar ports, including descriptions of the tooling it built to do it. I've done some reversing from binary but never with results this good.

This is a great idea. It could be a generic game porting skill to make the porting experience smoother. After my experience above. I have actually tried giving Claude binary files of games and it was able to reverse engineer to GoDot!

> I have actually tried giving Claude binary files of games and it was able to reverse engineer to GoDot!

Oh wow...I hadn't even thought of trying something like this. There are soooo many games from that era that I'd love to play again without the hackery of emulation.


Just grab any game, especially old rom files form the 80s (I tried TimePilot and Road Fighter from MSX era). The model extracted the graphics, the sounds and then built the whole game. It takes less than 1 hour to port a 16kb rom file, in one shot. You will get a playable game, but you need a few follow up prompts to get it back to its original form. Usually some sprites or sounds issues, but very easy to fix, if you have played the original game.

So by old ROM files, you mean still PC games or MSX, Sega, NES, SNES? And when you have them in place what do you specifically ask Claude? This is interesting and I'd love to see if it can get me porting old games with zero coding knowledge.

Yes, you can try your game out. I tried with old msx rom files and it worked.

Yeah, someone reverse engineered Red Alert 2 recently. Took Fable cranking for almost a month, but it's pretty amazing to see.

https://github.com/harshilmathur/openra2

this is the closest thing I could find. do you have a link to the reverse engineered RA2? I am surprised I do not get any recent repositories or forum posts about this, you'd think it would be more talked about


Dumb question this, but how do you load it into Claude? I’ve only seen APIs for text and image uploads. Not random files. Or maybe that’s a limitation of the SDK I’m using?

With Claude Code command line, you can point to any directory you have in your file system. In this case, Claude Code was working across three directories: the original game in assembly, the 2010 port in C++, and the new Godot port.

This is amazing. How long did that process take? Was it almost just a case of giving Claude the binary files and then prompting "please recreate this in Godot"?

I tried it with 16k roms. It does take less than 1 hour but if you want to tweak and make playable, you will need half day

I'm doing something similar for DOS SDL ports. Repo isn't public yet, but idea is to have a repo with a knowledge base, skills and patches agents can use when porting to avoid having to bootstrap everything from scratch.

Porting to DOS sounds like a great plan for retro games, to only have to do the port once and know there are stable emulators to run it in, as opposed to porting to some modern OS and then have to port the game again and again to keep up with API changes.

I'm assuming they mean SDL ports of DOS games. SDL is actually one of the most stable libraries that a native game can use. They make breaking changes to the API once every 10-15 years, and they write a compatibility layer that implements the previous version's API using the current version's API as a backend. This has happened twice so far, so there's an SDL 1.2 implementation that uses SDL 2 and an SDL 2 implementation that uses SDL 3. These can be chained together, so for example you can run an SDL 1.2 game from the early 2000s on a modern Linux system and it'll have native Wayland and Pipewire support thanks to using SDL 3 behind the scenes.

session stored on the drive would also contain all of the tool calls , right? unless op changed the default settings it would've expired though by now, since the run was in July

That's a good point. I looked at the git history for my analysis. What is the best way to analyze claude sessions history? I definitely check if I still have them stored. Thanks

sessions are stored in the home folder somewhere but are deleted after 30 days by default. it won't be stored in the bash history, at least it does not for me on macOS with zsh + omz

Of course is one does nightly, off site backups, all is well! A friendly note to all, always have backups!

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: