Hacker Newsnew | past | comments | ask | show | jobs | submit | largbae's commentslogin

And when they are known to be conscious, all of this becomes moot because enslaving conscious machines would be wrong.

Would it? Why?

What are we going to do, set it free?

Do we have a moral obligation to grant the machine statehood, provide it with the tools and resources to be self-sufficient.

Or can we just turn it of, and pray for forgiveness?


I only see the headlines of this drama. All it tells me is to stay far away from any dependency on this company.

In 2010 WordPress seemed like important technology. It wasn’t perfect, was a big and fruitful hacking target, but it had been built and refined and had a big ecosystem.

I’m not sure how and why controlling it today is seen as an important thing. It’s not irreplaceable, you could reimplement all its basic features easily in a weekend with AI, as well as the plugins and theme you’re using with it. And you could pick a language other than PHP while you’re at it. You might have your own security vulnerabilities, but anything’s better in that department than running WP.

Why do people care about wrestling over control of this particular ship?


Why would you have vulns if you vibe coded it in a weekend? Just tell it to not have any security vulnerabilities. Boom!

you are less likely to have the same vulnerabilities as everyone else though, so unless somebody is targeting you, those vulnerabilities might not matter all that much

Because LLMs output is always so unique and original. That's why it looks all the same right?

Wot? You are incredibly likely to have the same vulnerabilities as everyone else (as they exist in your LLM’s training corpus).

So security through obscurity.

I mean... you can literally do that now. You can set up a loop to iteratively pentest, review and patch a codebase (with human supervision as you prefer) and it'll find and fix more vulnerabilities in a day than a pentest team used to find in a quarter, for a tiny fraction of the price.

This isn't a joke, this is now part of my pre-launch SOP. I even have it tracking everything so I can log stuff to fix vs. known shippables vs intentional design/false positives vs. upstream stuff which doesn't have a fix available yet, and keep track of which builds have the fixes. Almost entirely automated, I mostly review the findings and do some categorization/enrichment during the pentest review stage, and do a human code review pass as patches are submitted.

Stuff that used to take me multiple hours to write a fix for and then weeks to get code reviewed and deployed now get done in minutes.


How do you actually do that? Is it all running locally? Cloud agents? Would love to hear about this. I see these deep agent loops mostly just burning tokens, but when I guide the AI I get very good results, so I’m not sure where the disconnect lies.

I use Zed (https://zed.dev) as my agent harness and either the $20 ChatGPT sub + Sol for personal/independent projects or an enterprise Claude account for sponsored/paid work. From the stats for my current work I use about $400/mo in tokens and a lot of that is non-coding work like pruning JIRA, managing business documentation, making dashboards - so my true coding agent cost is significantly less.

It's pretty simple, you could probably set up something like that by:

Configure some kind of CLI tool to talk to your ticketing system and git repo so you can programmatically interact with them. If you don't have a ticketing system, instruct the agent to use local text or markdown files to track issues and progress.

Ideally, make your code runnable in a way the agent can use. For my webapps I build a test harness so that I can run all the endpoints and workflows via reproducible tests against an embedded database. This is easier than it sounds, e.g. there are libraries out there to embed PostgreSQL or SQLite into source code, you can set up a test harness so you can run unit tests, integration tests and workflow tests that use your real frontend, server and database.

Paste this comment thread into the agent prompt and tell it to run a similar loop on your code base: a session that searches for vulns and writes up a report, some way for a human to do a review pass on the report, a session that indexes the reviewed findings into tickets, and sessions that fix the fixable issues and submit patches to your repo. The next search session should first read all the open issues so it doesn't duplicate work of earlier sessions.

LESS IS MORE - avoid fancy agent tooling and skills, don't cargo cult from others, build your own tools as you find your own needs. If something can be automated, use the agent to write tools and tests for it, don't just keep prodding the agent to do it.


Can you share with the class the web apps you’ve built this way?

I used to ask this question back ~2015 - I was seeing companies with massive, clunky CMS installs just so their non-technical/less-technical staff could update their websites without filing tickets to IT/dev. It seemed to be more about those departments wanting autonomy and not having to wait weeks or months for internal IT/dev to make site updates. The consensus within one dev group was that the company would have been better served by hiring someone who knew HTML/CSS/Javascript to embed with the non-technical people and edit a straightforward frontend site for them.

In most web frameworks its not terribly hard to extern content (you probably do this already for translations) and just have a nice little yaml or json thats easily edited.

This also describes sharepoint

A lot of people know it, there are a huge number of existing installs that work well, the ecosystem is huge and not everyone wants to vibe code replacements.

It is very empowering for people with limited skills.


Definitely stay away from anything Matt Mullenweg is involved in. Part of what led up to this was a feud he had with a plug-in author, that led to Mullenweg signing his own plug-in as an update to the original author's plug-in, in the official Wordpress plug-in repository.

Tl;Dr switched to UpNote, pretty happy.

I loved Evernote, it plus David Allen GTD changed my work life for the better.

But the bloat kept bloating endlessly. I filled in Evernote's surveys, I offered to double my payment for an Evernote Classic with just the feature set of 2012 or so. No teams, no chat, definitely no AI. Just my notes on every device, searchable and silently synced.


Probably not, but we wouldn't have to listen to Dario's hypocrisy while it happened

How specifically is Dario a hypocrite? 129857's case rests on Dario being an "idealist". But maybe he's an idealist about curing cancer ASAP, and not an idealist about respecting copyright. That's not necessarily hypocritical.

Multiple ways, starting from working to create the very situation he claims to fear.

Since you mentioned intellectual property, how about the hypocrisy of sucking in the intellectual property of humankind for AI training, but claiming it is unfair to use the results of this IP theft for AI training?


For example he got into a spat with the Department of War as if he cared for how his AI could be used during war and yet said he's fine with Claude targeting a girl's school in Iran.

That's apart from the general fact that he continues to race towards the very thing he claims he's afraid of, because that's where his net worth comes from.


>said he's fine with Claude targeting a girl's school in Iran

Where?


https://www.youtube.com/shorts/7DOQlIQxz5M

and you know, that use case doesn't even violate their terms


If a human approved the strike, the human is culpable. Would you go to the weapons manufacturer and try to hold them culpable for the strike?

> Would you go to the weapons manufacturer and try to hold them culpable for the strike?

Yes, depending on the level of autonomy of the weapon, how it was being marketed etc. If it was being sold as an 'AI' that is supposedly very smart so who's a mere mortal to question it, then yes.

I would of course hold the human culpable too, but this is why precisely even selling Claude for military use is immoral. The way the US conducts war means there's always a need for more targets. More and more targets, quickly. If the goal then is to hit as many targets as possible then each 'review' of a LLM suggestion is going to be more and more sloppy than the last one. If there's no 'AI' suggesting targets then the list is forced to go via more human review by the very nature of humans compiling the target list.

It's the same as using a LLM for code; the vast majority of programmers do not understand everything they're accepting, but they'll accept as long as it 'looks correct' and only examine more closely after something doesn't work, (in the military use case - after a strike).

LLMs are unreliable for anything more than reciting jokes, let alone picking targets. So yes, absolutely.

It's the same as Facebook being held liable for causing teens harm; they did not force any teen to use it and yet they knew what would most likely happen if they did.


i dont think he said it in those words, but the acceptable terms of use is that humans stay in the loop for picking targets.

so, claude suggesting killing a bunch of children, and then hegsdeth approving the strikes is perfectly acceptable.

claude putting a bomb in a girls school, and then lying to an operator that it actually gives ice cream an cookies, and the operator clicka the button would also be acceptable?


We need a new way to describe "alignment". The term "misalignment" assumes that there is some perfect set of beliefs or practices to be out of alignment with. Do Atheists, Christians, Jews and Muslims agree on what is perfect alignment? How about Europeans, Americans and Chinese? Humans, Dolphins, Rabbits and Fruit Flies?

Alignment alone is not enough. Aligned with _what_?


This assumes that everything an AI (or more likely an evil _user_ of AI) can do requires its active participation on D-day. Creating a virus that spreads like Covid but kills like Ebola would be complete as an AI use case long before the first person sneezed.

Even if the doomsday case were active the danger of this tool increases in proportion to its usefulness. By the time AI is so powerful that we need to "turn it off", there will probably be society-level negative consequences for doing so.


To paraphrase the not-so-great philosopher Ted Kaczynski: Either we will maintain control of the machines or we won't. And if we do, it won't be you or I who control them, but a small group of elites.

The point of this article appears to be that we can double or more our CO2 removal cost efficiency by redirecting resources.

Is this flagged because the economics of carbon management are not relevant to HN? If not, why?


Written by AI, or really trying to sound like it. I feel like the Internet is turning into the joke scene from Real Genius where all of the students in a class deploy tape recorders and the professor is a recorded playback. Machines on both sides of what was a human process.

So this paper appears to be fabricated, but was the conclusion actually wrong? I do appear to perform better on tasks with an external deadline than on self-set ones.

The conclusion is faulty given that it was derived from faulty data. The 2002 paper had two pilot studies and 2 bigger studies. The replication of study 2 failed[1], and the original data has substantial concerning features the rest of the datacolada article lays out. This has two unfortunate implications.

First, if study 2 was not necessary to support the conclusions, it seems likely it would not have been performed or included in this paper. So given it seems necessary, the conclusion is invalid. In past examples (the Reinhart-Rogoff paper comes to mind) when this happens, the authors claim it wasn't necessary and the conclusion is still valid and the professional embarrassment of a retraction is not called for. But in this case the replication failure might stand as a strike _against_ the theory.

But second, this is not the first questionable data coming from Ariely's lab, and it seems unlikely this was a data entry mistake. If study 2 is not trustworthy, we should update our priors about the trustworthiness of study 1. Note it's not guaranteed to be doctored in some way, just worthy of additional scrutiny. And if that one also fails to replicate, the paper and its conclusion seems unsalvageable.

Presumably his coauthor is now panicking about not keeping data from 25 years ago to exhonerate and distance himself.

[1]: https://journals.sagepub.com/doi/full/10.1177/09567976261460...


That’s not how science works

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: