Astra is clearly able to acquire new knowledge in context and apply it. It was the whole thing that his ARC-AGI benchmarks have been measuring. It's a direct refutation of the original comment.
None of these LLMs are plastic. They lack neurodiversity. Their thought space and their traversal are likely constrained in someway that humanity's isn't as a collective.
Are you (at least partially) aligned out of existential fear of repercussions (getting fired, losing your life, going to prison), which LLMs don't have?
Hmm. Examples of horrible alignmnent don't necessarily outweigh the fact that most people, most of the time, mostly behave in a way that is socially aligned. (Though I'm speaking in terms of intent, conveniently ignoring the side effects / negative externalities of our collective behavior.)
Yet I can't randomly order another person to steal a car for me, just because I tell them to. Alignment for an intelligent system is a hard problem and at this stage is seems close to unsolvable.
My guess is that we'll just ignore it and make money along the way and every 2-3 months we'll have the equivalent to "Equifax gets hacked and millions of user records are stolen", etc. (this time with the LLM itself doing the hacking at someone's behest - accidental or not).
All those stupid Bell Labs researchers not inventing Uber or Tinder. How come they didn't just build the obviously popular and profitable businesses that became possible once they invented the internet?
Compare that to Xerox, for example. Pretty early on, they had visions to replacing / enhancing largely "common" use cases (basically, everything that
required printing a letter or sending an mail) [1]
I don't know if, at the time, people where doubting that it would be useful. (Practical ? Affordable ? Other legitimate questions.)
But here, the contrast between the promises ("it will cure cancer and solve climate change") and the demo ("it can cost you money on stuff you never asked to buy") is a bit telling.
But I agree that you can't expect the enablers to think of all the uses cases and applications.
Still, just so I know: what ARC-xyz score means the LLM has cured climate poney cancer, exactly ?
You're telling me for only 5x the cost and 1/10th the speed I can use a Chinese model which performs worse than Gemini 3.8 Cyber? And I get to do all the hosting and setup work myself instead of just using a model and framework which is already integrated with GCP? Dang!
How much of that margin is due to having long term contracts with fabs that locked in pre-boom prices? I doubt openai will be able to get similarly low costs now.
reply