I think taxes like these actually worsen the incentive structure. If it’s effectively more expensive to do training and inference, labs will be incentivized even more to improve efficiency. Opaque recurrence seems like a great way to improve efficiency, and it comes at the cost of safety
Even more opaque than the models already are? I take that risk.
I also love how the labs are apparently screaming "Stop us! Please stop us!!!" at the top of their lungs while completely unable to escape their own incentive structure...
> I also love how the labs are apparently screaming "Stop us! Please stop us!!!" at the top of their lungs while completely unable to escape their own incentive structure...
It seems like you're implying there's some better alternative, but you're just describing a race-to-the-bottom. It is completely reasonable for every competitor to want an external coordinator (i.e. regulator) to break the pathological competitive dynamics.
Oh no a reranker which optimizes for pclick (bias to cheaper items) and profit (bid and margin)
I don’t think the forest is dark (obvious financial incentives) and I don’t think we’re pre-Google (as you said, you can literally still find what you want, USING GOOGLE, even)
Huge gains? What was the impact of this change, considering the only context that we have is that it was not important to the customer? Answer should be expressed in terms of cost, revenue, profit.
I don’t doubt the CTO was exercising poor judgement, but I do doubt that op was communicating his concerns in a way that was tailored to the audience
> would an ML optimization loop have discovered transformers?
I think that’s exactly the kind of problem this group is looking to solve. You make a compelling intuitive argument, but that’s not the same thing as a proof
So they’re watermarking requests according to your environment variables and maybe changing a string format if you’re in a certain time zone? Am I missing something here? Where’s the five alarm fire?
It is not checking that is the problem, it is sending obfuscated information about the user without disclosure. That is unacceptable in any context, let alone a tool that requires an unprecedented level of trust.
I'm trying not to be flippant (let me know if I failed), but most tools with an online connection send back a lot of information about you, with just about the same amount of disclosure (somewhere in the EULA it says they may).
Especially if you play any online games with ranked or PvP, they are likely doing a ton of work to prevent cheaters - and this information is necessarily going to be obfuscated, to delay the cheaters/hackers in their efforts to work around it.
From Anthropic's point of view, they have a class of "cheaters" they're trying to detect - people trying to distill their models. Those people are of course going to try to work around any detection, so you can't send that signal in clear-text where it is easily detected and blocked.
This isn't an online game we are talking about, it is a fundamentally different kind of application, non-deterministic and capable of manipulation. The bar is much higher. They should collect analytics if it is necessary to provide the service, but they should do so transparently. The solution to snare cheaters can't be worth the sacrifice of trust through non-transparent collection of user data. Although this particular information might not seem sensitive, if they are willing to use this method at all, they are probably willing to use it in other ways.
There's nothing wrong with the transparent collection of analytics. I expect any software that I run to tell me what they are sending at a bare minimum, and ideally give me a choice. This is the common and acceptable approach for a software company that deserves your trust. A willingness to cross that line with something small does not lend trust for something big, and there's nothing really comparable for the level of trust that this software requires.
This entire thread has lost its collective mind. Tracking time zone has to be up there with IP address and referer in the list of "trivial things every company in the entire world collects about you", and these things are entirely trackable without your consent or even knowledge. You're gonna have to turn off the Internet.
It seems like the confusion is about whether it is common practice to collect analytics vs whether it is common/acceptable practice to do so without disclosure and to cover your tracks. I recognize that virtually every cloud service collects info. Even something as trivial as IP address and region are not too trivial to include in the EULA.
What is the harm of that disclosure? If you are a company that wants to establish a firm basis of trust because you deliver a sensitive service, it is a necessity. I can't say with certainty that Anthropic doesn't disclose that they collect this info, but in any case, the way they've chosen to implement it does not lend credibility.
In your example, service providers may be allowed to collect IP address or cookies or referer which are either critical to how they provide service (IP and cookies) or part of the established norm (browser sending referer).
That doesn’t mean they are allowed to go beyond and scan user’s environment variables
Maybe you missed it in the post, they look for the value of an environment variable and send the obfuscated result. I'm guessing this particular variable isn't sensitive for most users, but the reality that they find it acceptable to snoop your environment and hide their tracks is a big red flag.
> The worst case of these were the few companies that set up token leaderboards, which is perhaps the dumbest way possible to encourage learning how to use LLMs well
My company does something dumber now. A leaderboard of how many lines of code you shipped, weighted by how complex they were (assigned by a heuristic). You can imagine the incentive this creates. I wish we just measured tokens
I’m confused: is it just markdown files in git? Or does the hybrid graph+semantic layer matter? If the latter is true, the title is just clickbait right
I tried semaglutide and while it was effective for losing weight, it made working out impossible (felt exhausted and very sick as soon as hr went up) and it made hangovers awful. Is retatrutide any different?
Had the same issue. With reta I have more energy than normal and workout even more than usual. I get great workouts despite eating around 1400 calories a day. I don't wanna act like it's without flaws but I gotta say it does it's job very well
reply