Some people posted about how to do cheaper better faster vision and it ended up in fable.
Some people posted about how to do cheaper better faster speculative decoding and it ended up in grok 4.5.
Details? Evidence? This is important and interesting if true, but if you just post the assertion without any substantiation, that's not very helpful.
There is a small short term penalty to your business if your tool is honest with customers, but the honesty pays for itself within a year
It seems like there are two unrelated ideas here, can you help bridge them?
(1) Posting new techniques online means those techniques get incorporated into the next gen of frontier LLM (or do you mean that the technique is in the training data so the next gen frontier LLM is merely aware of the technique?).
(2) Agents have dramatically different safe agent-hours, and this doesn't correspond with how "good" the models are.
those techniques get incorporated into the next gen of frontier LLM (or do you mean that the technique is in the training data so the next gen frontier LLM is merely aware of the technique?).
I mean the technique will be directly used.
Agents have dramatically different safe agent-hours
Ah my point was that this um safety-metric, like many others, is cheap and easy to get. The unobserved extreme variance on break-your-computerness proves that it is attainable. I'm not sure how to explain this. It's like half the sports cars explode, and you can prevent it with a couple little gaskets, and nobody noticed. This is a point of extreme leverage. One talented person can tilt the scales between "sudo delete humanity" and "askuser would you like to delete humanity"
Indeed it probably will come down to the presence or absence of that person.
What a time to be alive!
Some people posted about how to do cheaper better faster vision and it ended up in fable.
Some people posted about how to do cheaper better faster speculative decoding and it ended up in grok 4.5.
There's currently massive differences between models in how long you can leave it unattended running on host with sudo without things going haywire
I don't have a bar chart, but opus tends to break my hosts within 2 agent-hours, gemini 3 pro within 12 agent-hours, and gpt 5.5 within 24 agent-hours. Then they use all ram or change ssh perms or delete the data or whatever. I ran qwen 3.5 122b on bare host for 50 agent-weeks (50 parallel for one week) with no observed ill effects!
All the labs and all their customers want "the thing i asked for actually got done" and they all want "the model's summary reflects reality" and so on.
What an opportunity! You can just post a method for "how to make ai tell truth" or "how to minimize side effects" and it will probably end up in the next frontier models