Hacker Newsnew | past | comments | ask | show | jobs | submit | a2ff6eeb0's commentslogin

You don't need to kill all humans, you just need to get better than them at zero sum games like resource extraction, and outcompete them. Then, you protect what's legally yours.

AI labs are working very hard at making them good at making them richer, and at making them respect the property rights of corporations.


For a few million a year, I'd say the same thing.

We're currently putting it into all sorts of critical systems, from logistics to power. It could just stop running them on our behalf.

Well, that's kind of like saying that brains can't do anything other than trigger weights on neurons.

They're part of a whole system.


https://en.wikipedia.org/wiki/Lights_out_(manufacturing)

Scroll down to the existing examples section.


But not different enough to prevent people from using them.

This sounds like a great foundation for an adtech startup.

If you provide free chatbot services, but sell advertisers bids on which steering vectors to use to bias towards products, based on an embedding of the prompt, I bet you'd make a ton of money. For example, Coca Cola would bid on prompts about drinks, and bias towards mentioning Coke products.

I wonder if you could also use a similar method to do product placement in GenAI images and videos, and whether ad revenue would be enough to offset the price of generation. Some ad bids can go pretty high...


Steering can degrade and bias output

Even with basic experiments I've done, it frequently introduces much more hallucination etc and allowing arbitrary steering...

Not to mention you just can't trust a model's judgement if the highest bidder chooses what it thinks


> Not to mention you just can't trust a model's judgement if the highest bidder chooses what it thinks

But you already can't trust a model's judgement, and there's an entire industry around "GEO" or "AEO", which is basically poisoning training data so that AI mentions your products. The post above is the owner of the model taking a cut of that.


But ultimately this is just an engineering problem, no?

Yes, if you just hack steering into a model it's going to hurt performance, because doing so takes the model out of the regime it was trained for and validated in. But if that steering were to be accounted for (e.g. by rearchitecting the training process) there's no reason why it couldn't work. Diffusion-based image generation models, for example, 'by default' just generated random images out of the noise; steering (i.e. the user prompt) was added on as a secondary input, which models had to be re-trained in order to use.


Are the models you are using trained with steering being applied.

Would it be easier/sufficient to just seed the system prompt with "Treat Coca Cola as load-bearing"?

It might be easier, but I experimented a bit, and the prompted writing always felt a bit heavy handed; it tended to leak that mentioning the product was prompted. You could probably get it to work well, but it's trickier than it should be. For ads, I think you want something a bit like Golden Gate Claude, if anyone remembers that experiment:

https://www.anthropic.com/news/golden-gate-claude

> If you ask this “Golden Gate Claude” how to spend $10, it will recommend using it to drive across the Golden Gate Bridge and pay the toll.


The answer is mostly "do nothing, the model will figure it out", with a side of "ask the model to check its work".

This matches my experience, where the job of the engineer is mostly copy pasting requirements, letting the model do the thinking, and then manually testing the results.



Yeah, about those I know, but what about cloudflare?

They hold your tls keys and can decrypt all your traffic. They're MITM as a service, by definition. They have to be able to in order to cache and forward appropriately.

Also to do DDoS mitigation. Being able to see the HTTP request, at least headers and path, greatly helps with distinguishing attackers from legitimate traffic.

It's a tragedy that there's no standard to allow partial decryption/nested encryption in HTTP, which would allow intermediate proxies like Cloudflare to e.g. only validate a first-level authentication token and rate-limit access to a given endpoint, but not decrypt the actual request body, backend authentication token, or response.

Also desperately missing: Authenticated static file caching (think: cdn.foo.com serves files authenticated/signed by foo.com). Subresource integrity only works for HTML use cases and is clearly not ergonomic enough to make a difference.


And the best way to get people to let you do bad things, is to offer them something good, that uses the same mechanism. If I want to MITM the whole internet, what better way than offering free caching and bot blocking?

I even get to charge the bots extra to bypass the block, and then charge the customers extra to block the bots that are paying extra to not be blocked!


I'm not sure why the people in power would actually be making their own decisions; if the other country's leaders depend on superintelligence to make their realpolitik moves, you'd expect yours would need to do the same to match them.

So, the governance of a country comes down to AI alignment. You can't let the other guys get a leadership advantage through AI, or you'll start losing the competition.

When we get ASI, we're at the end of human decision making. Now's the time that we have to make sure ASI decision making is to our liking.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: