At what point would you, as a chimpanzee, have been worried about humans potentially unseating you and threatening you to the point of one day being an endangered species on the brink of extinction?
By the point you would have been worried, would it have been too late?
Problem is this argument can be leveraged to wipe out any living or non-living thing whos numbers pose a potential threat. Other religious groups, races, even sufficiently different cultures.
Who killed the Neanderthals? Were sapiens actually smarter or were they just less accepting of those different than them?
I think this is clear evidence that AI models are now at the far frontier of mathematics innovation and discovery and exceed human limits.
This specific problem having had a $1 million bounty on its head and still remaining unsolved for 26 years after the bounty was placed is pretty clear evidence that many of the world's best human mathematicians would have solved this problem if they could have, and none were able to until LLMs came along.
Hard to claim at this point that LLMs aren't capable of novel STEM creativity and genius to a degree that will soon far surpass that of humans.
If anyone has counterpoints to this I'd love to hear them!
Not a counterpoint per se, but I burned $50k recently on a much more modest math problem (result already known, just thought I had a sketch of a more interesting proof), and the LLM thought it had proved it within those bounds but had instead subtly fucked up the Lean definition. Take from that what you will.
Not to mention, it's still very much up in the air whether the model derived the answer of its own accord or sniped the important details from the researchers it was spying on.
Mostly not my money, tech makes one fabulously wealthy, I care more about math than fast cars or whatever (even as a pizza driver you can afford a fancy car if that's your primary motive), etc.
> Already solved
That's a very interesting insight into mathematics. It's absolutely just as interesting to prove that certain proof techniques are or aren't possible as it is to actually prove the main result, sometimes moreso, especially if they have any chance of improving attacks at other problems.
To be fair, I think it's still an open question about how far it might surpass human capabilities.
I think it's clear that its speed of development will be significantly faster, but it's technically not proven that the frontier and problems don't themselves become increasingly difficult faster than any acceleration in intelligence past the point of human training, data and existing knowledge.
Should this be the case, we would see a rapid broadening of development, and a slow advance in the frontier in such a way that might surpass the collective capabilities of people, but not by very far.
Fields that allow verification, like math, will far surpass human level because they don't need human data for training. It's exactly the same as with Chess
My point is that you are probably correct but it is not proven.
If the field that requires verification has an increase in difficulty which exceeds the increase in intelligence, than the capabilities might surpass human months but not by far.
For example, we saw a rapid decrease in space exploration progress after the discovery of the Moon because it is much much more difficult to get to Mars compared to the Moon. What we don't know is what the frontier looks like. It may have a gradual ramp up in difficulty, but it might also involve some sort of cliffs which even artificial intelligence is unable to climb.
I think the biggest hurdle remaining is that all these landmark results are generally counter-examples.
Proving something in the affirmative often requires the creation of an entire new sub-field of math, or new tools. Think of Fermat's Last Theorem or something like that.
These results, while impressive, are clever constructions using existing techniques. It isn't clear that AIs can build new machinery like this. But if/when they can, yeah it is probably game over.
When you read the detail the compute they are throwing at it is incredible, tens of thousands of agents with different groups competing.
It's not like a single Gauss as you imply, "just" many, many mathematicians working tirelessly in a completely ego-less way, guided by other agents and ultimately humans, built - allegedly - on recent human insights.
Stunning, undoubtedly, but this is a "brilliant autistic herd" result, not that of a singular mind.
> this is a "brilliant autistic herd" result, not that of a singular mind.
I slightly disagree. A single LLM is equally 'mindless' as a herd of them. As anyone will tell you they "simply predict the most likely next token," yet, complex solutions to difficult problems arise from them.
Many people have said that the architecture of LLMs will need to change for true ASI. I think that the herd of tens of thousands of agents can be seen as one such potential architectural extension. Whether or not a herd or a single LLM is used for a result like this is irrelevant.
To be clear, I think the orchestration of thousands of LLMs in their current form, even with ever increasing intelligence, is not the form ASI will take. There is still a major architectural breakthrough to come, in my limited, ignorant opinion.
Sure, even a 20% chance at 1 million payday after 5-6 years of fulltime work on a project with zero practical application doesn't touch the, say, 200k/year guaranteed our best mathematicians would have to forgo to devote their intellect to the problem.
Are these mutually exclusive? Why would you have to forego that salary to work on this problem? This is one of the most prestigious and meaningful problems in all of mathematics, which is why it has such a high prize amount attached to it - why would a university not support a mathematician working on such a prestigious and important problem in lieu of something else?
i wonder how many tokens it takes to run 10,000 agents? One could argue this is simply a problem of appropriations. I find myself wondering if a corporation could spend $5M on mathmeticians and arrive at the same end result.
>If anyone has counterpoints to this I'd love to hear them!
Sure. A proof without an unknown amount of human steering (and/or stolen research) would be an unquestionable achievement.
To this day there's zero (0) evidence of any result by an LLM alone (maybe I'm wrong). If I just prompt ChatGPT right now with "give me a proof of the Riemann Hypothesis" and this thing delivers, I'm sold. But anything close to "yeah ChatGPT proved X with 5 years of 24/7 work with 10x Terrence Tao level geniuses" it really doesn't cut it.
Or why's there's no new branch of mathematics invented by AI? That'd be indubitably _novel_ and _creative_. But to my knowledge (and I'm eager to be educated) there's nothing like that. What are the HARD examples of novelty, creativity and genius you claim? For how people like you talk about AI I'd expect idk, a unified theory on fundamental physics, or a novel engineering solution for material science and nuclear fusion, or at least improve itself to not need a bazillion GPUs to emulate a 20 watts wetware. Sure it would infinitely easier to make OpenAI literally print money with any of the thousand problems easier to solve with such amazing intelligence than the NSE problem right? Honest question
> I'd expect idk, a unified theory on fundamental physics, or a novel engineering solution for material science and nuclear fusion, or at least improve itself to not need a bazillion GPUs to emulate a 20 watts wetware.
I think the counterpoint here is simply to look at what was being achieved with LLMs one year ago versus today, and extrapolate that trend. Sure, there may not be examples of what you've asked for yet, but Astra is literally a couple of months old, the model that solved Navier-Stokes is less than two weeks old. It appears that we're seeing the hockey stick that only the most bullish thought was possible.
Parent comment is claiming creativity and genius beyond human experts, so why not ask for a fully unassisted AI novel result? Having access to the entire corpus of human knowledge, what else such amazing entity would require to solve a hard problem by its own?
Any result of such kind from an AI alone would be enough to refute my argument, yet you don't present any.
>Also there are proofs where the only human steering was "keep going".
>Throughout this process, Jarred's input was mostly limited to sending Claude messages of encouragement (mostly variants of “keep going” or “believe in yourself”).2 This seems to have helped Claude overcome some initial skepticism that it could make meaningful progress.
> Drawing on extensive prior research by mathematicians over the past decades, it
> [Claude] has increased this bound [for the fraction of zeros of the Riemann zeta
> function that satisfy the Riemann hypothesis] from 41.6% to 67.2%. Claude also
> produced a formally verifiable proof of its result.
How is a formally verifiable proof not a proof? You're making literally no sense.
The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage they show for Opus 5 which would similarly be much higher.
Regardless, the result is still valid as the original benchmark harness is definitely unreasonably handicapped, and if a harness alone can help the LLM saturate the benchmark with a near perfect score then the combination of the two must still be effectively AGI in the sense of passing the most famous benchmark designed specifically to measure AGI progress, after multiple iterations of progressively making it harder.
I think it is fair to say that this is probably effectively AGI if the benchmarks are remotely accurate - even with Fable, I've been at the point personally where I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks. If Astra's this much better than Fable, I'm ready to call AGI here.
For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.
When we released ARC 3, I got asked, "when do you think a frontier model will saturate it?", and I answered "in about a year, though it depends on how much it gets explicitly targeted"
That was 6 months ago, so the progress that Astra represents happened about 2x faster than I anticipated. I think the speed of progress will surprise a lot of people, and what the new models can do will challenge the views of AI that people developed by using prior generations of models.
I feel like AGI's definition got watered down, and these tests do not cover the original definition, what is your definition and thoughts on aligning with what all of us understood from the original claim?
I feel like this test is just helping someone like Sam Altman pretend like he implemented AGI as originally pitched for an IPO when in fact, he has not. Shameful.
> AGI is essentially the equivalent of a median human that could be hired as a remote co-worker... capable of performing any task that one would be satisfied with a remote colleague doing via a computer.
I can call a coworker right now and have a real time conversation with them without feeling like I’m talking to a frustrating machine. Most of them, anyway. But I guess that’s “moving the goalposts”.
> I can call a coworker right now and have a real time conversation with them without feeling like I’m talking to a frustrating machine.
These machines were designed with the express purpose of being condescendingly sycophant to the point they hallucinate just to state "you are absolutely right".
No wonder some people even find these chatbots to be wife material.
And it disagrees with me a lot, granted this is in the webchat which I don't use for programming but for general usage seems to work, it isn't sycophantic
I don't have those system prompts, but gemini pushes back on its own quite a bit. It'll tell me when I wrong, or when there are better options to consider. They won't be exhaustive, but good enough for 95% of my queries, so it's also my go-to web chat.
Random humans don't have a significant cross section of human knowledge available in real-time, although many like to pretend they do, especially in internet comments :P Being able to compete with the capabilities of a median human would be an absolutely world changing achievement.
What do you mean? Virtually all humans have access to the internet. That's literally "a significant cross section of human knowledge available in real-time".
Oh, the median human can't process that in realtime, you say? Looks like they can't compete with the capabilities of the frontier AI then.
Yep - I like to phrase it as "AI is better at most tasks than most people". AI will still be beat at experts at specific tasks, but in general I find it to be better than me at the areas where I have no expertise. It's the ultimate generalist.
Have you? It’s a dumb model that gets a lot wrong. Codex voice is IMO only good for when I cannot dictate into the good models. The delay introduced by letting it work with a good model kills it for me.
Would I trust to let an AI, with zero human input or oversight, to diagnose, come up with treatment plan, and ultimately operate on my l5/s1 disc that's been bugging me for the better part of my adult life?
Would I take a novel drug "discovered" by AI (I mean entirely by AI, no human input, remember we are talking AGI) that promises to cure some chronic neurological disorder?
In both of those cases, they are the biggest hell-no's I can emphatically say.
Until I can say hell yes to that question, we aren't close.
Preempting those who say "Well your doctor/drug companies are probably mostly using/going to be using AI to do that" -- not what we are talking about here, and in both cases, not AGI (and I would probably find a new doctor)
Yes, but it means general the way humans are general. Clearly being a general intelligence shouldn’t require being any better at any individual task than the average human, or even the bottom decile of humans.
Most short-term learning/adaptation is already handled in-context. Modern context windows can hold several books worth of text - plenty for most tasks. Everybody is already using it to adapt models to their projects through skills/instructions/guides etc. ps. I often say that after glossary-skill next must have one is update-skill-skill that threats all .md files as live documents.
Persistent weight adaptation also happens just not in real time - sessions are captured, analyzed, transformed into training data, fed into SFT/RL environments and later contribute to model updates. Takes a bit of time for the whole loop but you can't say it's not present.
There's nothing fundamentally preventing real-time weight updates, ie. LoRA-style online adaptation would be one obvious approach. It's just generally not worth doing at scale. Updating a shared model centrally gives much better data efficiency, batching, evaluation, control etc. than continuously training a separate set of weights for every user/session.
There is some work happening on narrowing that gap, for example Mistral has been pushing efficient LoRA-based customization, continuous pretraining, model adaptation etc.
I also did play a bit with activation steering – it's super cool where you extract profile for some concepts (emotional in my case) and you have effectively toggles to control "brightness/contrast" those areas (enhancing or suppressing those activation regions from profile) injecting to the model those concepts (emotions in my case) – you can do it in real time and it's fun thing to play with.
My Claude admits mistakes and then fixes them on its own all the time.
Often even without my input: "(thinking..) Oh I discovered that I misjudged XYX, let me fix that.. (thinking) (executing scripts) Okay I corrected my mistake, I had accidentily ABC."
I may miss some context because the GP’s link has a paywall. But Altman said, "Let’s say we make an AI that’s really good". What is that supposed to mean? Really good relative to what? Current models are "really good" in many ways but nowhere near AGI. Really good compared to an average human? At everything? We’re talking about AGI, so "it’s a really good programmer/coworker/whatever" is a necessary but nowhere near sufficient condition, obviously. But given the constraints of LLMs, we can cut them some slack and only demand they be human-equivalent at digital tasks rather than walking and cooking. Still, being "really good" at all that seems to me really difficult to measure. But it’s a sufficient but not necessary condition anyway, AGI just means equivalent to human, not equivalent to a really smart human. So I really do wonder what Altman meant there if anything.
"at everything"? Surely you know a friend or two who is not good at almost anything – but you wouldn't hesitate to say that he possesses general intelligence.
What if it can be Einstein, but can't draw a Pelican, write a solid college-level essay, or fold clothes?
The ability to do a ton of book learning in training, and pull in tons of related context at once, is superhuman in some ways, but lags a lot in others.
> What if it can be Einstein, but can’t draw a Pelican, write a solid college-level essay, or fold clothes?
Then it’s an expert system.
Stephen Hawking wasn’t very good at folding clothes.
The ‘General’ part of the term ‘AGI’ seems like a trap to me, because there will always be new workflows to master. Can Astra one-shot level completion on some yet-to-be-released video game? If no, does that mean it’s not yet ‘Generally’ intelligent?
You won’t get pure ‘general’ intelligence until you find Einstein’s hidden variables and load the state of the entire universe into context.
Meanwhile, building a series of expert systems targeting specific valuable workflows is useful today and seems like it’ll continue to scale to cover huge swathes of economically valuable workflows.
I think that’s the more interesting thing to be measuring. The surface area of useful economic workflows that can be addressed with expert systems built with today’s tech.
Hitting some ‘Artificial Expert Intelligence’ coverage threshold on economically valuable workflows is what will matter for humans well before pure ‘general’ intelligence.
The only important part of 'general' is the ability to learn from experiential data and update your own model. That's what leads to general capability. Humans can't oneshot any task natively, but we can practice for a while until we uncover often novel methods of accomplishing something.
Therefore: the current transformer architecture is fundamentally incapable of AGI because the models have no mutable long-term memory.
You only have weights (large immutable memory), or context (small mutable memory).
Humans have mutable long-term memory: I can learn a new skill, adapt an old skill to new information, or learn new knowledge today that I couldn't perform/didn't know yesterday. I don't have a training cutoff.
Context engineering is an attempt to paper over this limitation. You can get really far with context engineering and huge models, but you will never get to AGI because there are many tasks where humans' mutable long-term memory outperforms.
For example, a human can invent a new musical instrument and then learn how to play the instrument they just invented. That's inference (inventing an instrument) leading to training (neuroplasticity). Humans have the ability to train our NNs with considerably fewer training samples. Everything that you can do with transformers is in one causal direction: training -> inference.
So if we take a huge with enough compute (CPUs, b200s, petabytes of SSDs), we install on it both the Astra, and the toolsuite to incorporate new sensory inputs (threads/sessions), camera, microphone, temp sensors, the lot, into a new version of the model. This model is then swapped for the old model, or traffic slowly brought over, or even adjusting weights in place.
Then my hypothesis is that thing as a whole could achieve AGI.
This feels like a very close approximation on how we humans evolve our brain. By encountering new experiences/sensations, classifying them as negative or positive to us, filling it away in neurons. Or by training motor skills etc. In the end we get more connections between neurons in our brain and we are capable of more.
Bingo, LLM architecture just does not lend itself to becoming AGI. They can get really good, sure, but they will always struggle with novel input and scenarios.
The more training data that is shoved in to them, the more they'll seem to solve novel situations, but in reality it'll be things that exist in the training data.
Aka Star Trek hologram characters aren't sentient, and actually anyone who things droids in Star Wars can think of a weirdo. C3-PO just kept running out of context and trying to revert to it's system prompt.
> Can Astra one-shot level completion on some yet-to-be-released video game? If no, does that mean it’s not yet ‘Generally’ intelligent?
If a model can't learn on their own to play some new game just as well as humans do, it's not AGI.
It's okay if they would take some hours or days of learning (like humans might), but if they can't do it at all during their normal operation, that's not general intelligence
> You won’t get pure ‘general’ intelligence until you find Einstein’s hidden variables and load the state of the entire universe into context.
But humans have general intelligence. AGI is about matching human ability, and we know this is possible in principle because brains exist
And the only reason LLMs can't write essays indistinguishable from human output is because they aren't RLHF'ed to write like humans.
Folding clothes isn't an LLM's job but if you were to insist, they could certainly do it, as any number of videos from robotics labs will attest. That particular future is already here but definitely not evenly-distributed.
Adding sibling comments, I think some people may be overestimating how well the median human can draw a pelican, or create an SVG of a pelican (depending if we’re comparing to an image generation model, or SVG generation).
Most people can't draw a bicycle. There was an artist 10 years ago that asked people to sketch a bike, and then turned these sketches into 3D renders - quite funny.
I can't draw a pelican. Literally my only point of reference would be AI pelican drawings from the test. Otherwise I wouldn't know how to draw one at all.
I would be able to draw an accurate bicycle, but I'm an outlier on that. Most people could not draw one [1].
Can definitely write college level essays and have for a while. The jobs is that when LLMs first started getting popular, but aren’t quite common professors were that some of the worst students in class started writing the best essays. Now everyone complains because they can detect the slop, but most human writing is so bad. But the really good human writing is still much better.
I would maybe argue that Einstein was the most LLM-like of great thinkers.
A lot of his great discoveries were mostly that he was very knowledgeable about the bleeding edge research in a number of disparate areas, and was able to have the aha moment where he could make the connections for how to integrate them.
A lot of other thinkers who created new fields from scratch are probably way harder for an LLM to crack.
That is very aligned with an LLMs ability to have superhuman knowledge in wide areas.
That’s like that scene in the I, Robot movie when Will Smith’s character is asking the robot “can you turn an emtpy canvas into a work of art, or compose a symphony?” and the robot replies “can you?”.
I think they mean improve our understanding of physics with new theoretical results or paradigms. Like if it’s 1899, would Astra develop General and Special relativity on its own?
This might just not be possible at the current time. In 1899, there was an "experimental overhang" in physics -- results that could not be explained theoretically (Michelson-Morley, but also lots and lots of empirical material/spectroscopic properties that we could today calculate using quantum mechanics). The big problem in theoretical physics today is that unifying general relativity and quantum theory has no experimental results you could get at our technological level.
I think you make a fair point, but also remember: new paradigms don't necessarily require confusing / contradictory observations. You could have the simple idea of "what if gravity is an inertial force?" at any period in time and work out the mathematics of this. It would make theoretical predictions which could then be falsified, but then again, who would take it seriously enough to test it if it was maybe say 1850 and not 1899.
A better example is maybe Maxwell's laws. Maxwell wasn't inventing a theory to try and explain confusing results, he was unifying a chaotic, empirical laws from existing experiments. That may be a cleaner example. That knowledge compression into satisfying theoretical framework is likely what is attractive.
You can potentially ask the same thing about like you say -- general relativity / quantum gravity but also likely plenty of other areas that may be like this today. Again going outside my particular area of expertise: standard model physics is in a large important sense empirical; lots of values and numbers that are simply unmotivated by theory or where we don't have a good way to make a principled theoretical choice. That could be a place where these models are able to help.
But right now: I doubt it. This is what everyone is working furiously on right now. How do you close a "science" verification loop? In principle this should be easy right: you have ideation (exploration, sampling with ~high temperature maybe as an analogue) and you have verification (which of these ideas are good) which amounts to rejection sampling in idea space. You have to have a sampler that is good at picking _good_ ideas for efficiency sake and you need a relatively fast and reliable verification step of "is this idea good and worth continuing to explore". But I may oversimplify
This is as good a time as any to note that we might be closing in on a new conceptual revolution in our own time as it relates to holography and an information centric approach to spacetime. Obviously It's the furthest possible thing from a guarantee, but it has much of the enthusiasm and motivation that string theory had previously enjoyed in prior decades.
So it could be a natural experiment for whether AI can contribute to novel physics. Specifically, there's a big question about weather. Something like our informational understanding of black holes where information inside it is equivalent to information on its boundary (which I'm sure I'm not saying correctly), might be generalized to regular space-time. More people should be freaking out with excitement about this and perhaps it's something to which AI can contribute.
Honestly I don't think I have single great article, though some Quanta ones are ok, and the Wikipedia article is okay.
The best thing I can recommend is what I did, which is ask Claude about the significance of (1) quantum computing error correction, and (2) error correction in black hole holography and research convergence between the two.
I have no idea what you're hoping the contribution would be. The AdS/CFT correspondence is 29 years old by now and it doesn't seem to apply to our spacetime, where the cosmological constant seems to be positive rather than negative. There are some puzzling consequences of the holographic principle in that scenario as well (https://arxiv.org/pdf/hep-th/0208013), but the linked articles don't talk about them?
Quanta articles are written for people with no background whatsoever, which makes them impenetrable if you have a bit of background and are trying to figure out what they're about. I don't know how good Claude is compared to that -- whenever I try asking any LLM about something I don't understand, it produces a wall of text, I have no idea whether it's correct or relevant, and I look for a textbook or review paper instead.
The wiki article has a section on 'Energy, matter, and information equivalence', the first Quanta article is almost entirely about the 'deep connection between quantum error correction and the nature of space, time and gravity' and about bringing the same information centric approach from AdS to our spacetime. The second Quanta article is explicitly about about bringing a holographic approach to our non AdS spacetime and cites an Ed Witten paper as the cornerstone of that approach (which one perhaps overly excited MIT physicist describes as 'revolutionary').
AdS is 'old' but the articles aren't suggesting it is new, and our spacetime is not AdS and the articles don't suggest otherwise. The point is that there's a search for a way to fit the holographic approach to our spacetime that's inspired by how AdS helps make sense of black holes. Quanta writing being directed at a lay audience ought to be a good thing, not a bad thing and they do link to papers if that's your jam.
Whether or not your LLM of choice produces indecipherable walls of text, and whether it ties those to sufficiently satisfying citations, I think is just a matter of how you go about the prompting.
You're right, I'm prejudiced against Quanta (often IMO they look for a clean narrative to the point of misleading and/or rely too much on metaphors) and was probably too harsh here. Sorry!
That said, I don't like them because their articles never leave me feeling like I understood something. They never go in an order of simple to complex and constantly try to hook you. (A positive example to contrast would be 3blue1brown, who manages to both hook you and make you understand, even with rather little background.)
Could you share a link to your Claude conversation? If this is a prompting issue, I would be interested in seeing what's possible. Thanks!
There was symbolic AI programs in the 1980’s that “discovered” Kepler’s laws and the resulting solar system model from just tycho brache’s astronomical observations. That was the the very first “new physics” ever.
Do you mean after? People do this!! But I think it’s a bit different. It won’t be apples to apples because the data volume I think is just so much different. Maybe there are good experiments for something like this.
For it to be like a human it wouldn't just need to solve existing phsyics problems, it would need to push the field forward and introduce new paradigms.
My comment wasn't very long, yet you somehow still ignored the main part, "and introduce new paradigms". The point is whether it can do everything humans can, entirely new theoretical frameworks and ideas, such as string theory or dark matter, are not coming out of AI at the moment.
AI could discover candidate novel physics without autonomously operating new physical experiments, and humans or instruments can later independently validate the result. This is analogous to how Einstein developed theories whose predictions were confirmed by experiments and observations only years or decades later.
Do we have any examples of an current day AI system introducing a novel concept or perspective. We've got plenty of counterexamples discovered and some theorems proven, but afaik nothing analogous to a new definition.
It could be a good theoretical physicist. Actually it could be a good experimental physicist as well since senior experimental physicists use grad students for the manual labor.
ARC-AGI was never "if this benchmark is saturated, we're at AGI". It was always about crafting adversarial tests that humans are good at, but current AIs are bad at. Point out the gap, get AI teams to attack them.
In practical terms? They usually get solved with a bigger badder LLM. "New ideas are needed?" Nah - ten times the params, ten times the test time compute.
ARC-AGI-3 was more of a failure in that regard than -1 or -2, because even on day 0, an off the shelf LLM with a harness could get 50%+. And messing with evals by forbidding "LLM with a harness" from scoring? Yeah no, that was just bad.
If it makes you feel any better the curmudgeon who drives the ARC-AGI tests feels the same way, that's why we're on 3 and I'm sure we'll see 4. Also, we can all avoid calling it "moving the goalposts" so no one feels talked down to.
If it can replace a worker but does too much work to be checked routinely by a human, and bears no real responsibility for its actions, well... it's really just a way to jack up the value of the settlement the company using it gets to pay out when it does something that causes a lawsuit.
If OpenAI had just simply stuck to making "good enough" models that were open sourced (like they promised they would be when starting out) and could be used to augment a human doing a task - a human that could be given actual consequences for messing up - they wouldn't have burned all of this money trying to reach this nebulous definition of AGI. Hell, "good enough" is what many open-source models are, and that's what terrifies Altman.
What do you mean by bearing no real responsibility for its actions? If I use a model to accomplish a task and it fails, I use something else to try to accomplish the task. If it's my responsibility to complete the task, it can't be the model's responsibility unless I've agreed to some sort of guarantee from the provider.
If the provider says "the model will always be right or your money back" then the provider has got responsibility. If there's no guarantee, there's no responsibility on their part, just on the person whose job it is to try and solve a problem with the model.
The "dream" that these labs are mostly selling is the ability for capital to subscribe to their AI for cheaper than it costs a human to do some task. Not to have a "human + AI hybrid where the human is responsible". It's what the whole AGI valuation is based off of, in that scenario, with no human oversight, the agent they lease has to be responsible for the task?
It's not really a dream though? You can subscribe to their AI right now and complete many tasks for cheaper than it would cost to pay a human to do those tasks. Another person can do the same thing but also keep a human in the loop. You and the other person may compete in the market for whatever your product or service is, and you both might do very well or one of you might do better than the other because of a whole host of different reasons. Nothing in the process of developing and selling access to a more advanced LLM requires the customers to do away with human labor, nor does it require them to offer the LLM service with a guarantee that it will never make any mistakes. So far the LLMs have always made lots of mistakes and the companies sure keep making a lot of money.
Right now I think it's still a dream for a few reasons, not the least of which is that all of this goes to shite if you have mass unemployment in a country with more firearms than people legally allowed to own them, a plurality of the population that treats wealth as an indication of personal virtue, and an elite that more-or-less refuses to offer any further evolution of the social safety net past what it was in 1970.
> Nothing in the process of developing and selling access to a more advanced LLM requires the customers to do away with human labor,
That's the hook, though. You acknowledge this yourself:
> You can subscribe to their AI right now and complete many tasks for cheaper than it would cost to pay a human to do those tasks
Businesses exist primarily to make money. It's an iron-clad rule that one must spend money to make money. If they have to spend less on humans to make the same amount of money, they'll do it. Furthermore, AI providers (especially hyper-scaling frontier model providers like Anthropic and OpenAI with insane operating costs) have every incentive to keep the price of their service as close as possible to the cost of the human. Ideally for them, you replace the human that cost $100,000.00 to employ by paying for a subscription that costs $99,999.99 while making the same revenues.
> So far the LLMs have always made lots of mistakes and the companies sure keep making a lot of money.
The clients, sort of. I know my team's velocity has increased. You can screw up with LLMs, like you say. OpenAI and Anthropic? lol no, they're massive furnaces for money, and will be until they can charge that $99,999.99.
> What do you mean by bearing no real responsibility for its actions?
If you give an "intelligent agent" offered by one of these model providers a task of updating the content of your website, and it updates it with inappropriate adult content, who incurs the cost of the machine's error? The model provider generally does not.
It it makes a mistake and deletes your website from AWS, who is responsible?
If it targets another website because it decides that it is "part" of your website and attempts to break into it, who is responsible?
In all of these cases, it's you. It would be the same if you downloaded an open model, ran it locally, and it happened to make the same catastrophic mistakes. The consequence to the provider is that if they offer a product that does these things, people don't buy the product.
In general, the person whose job it is to provide the company with a working, non-adult website and not hack into other websites is the one who would receive consequences for failing to meet those expectations.
The problem I see here is that ultimately, you'll have capital wanting to replace workers like others have said, and have someone roughly equivalent to a manager or vice president driving teams of agents to achieve business outcomes.
These tools can push out more results than a human can hope to evaluate in a business-sensitive, or even realistic, amount of time. You have to take it at its word that it did things right, and there's no real fear of failure or consequence on the behalf of the agent.
> If you give an "intelligent agent" offered by one of these model providers a task of updating the content of your website, and it updates it with inappropriate adult content, who incurs the cost of the machine's error? The model provider generally does not.
Something in your prompt led it to do that, is alex0015's point. The statistical odds of these frontier models screwing up to that extent are so impossibly low that it would almost have to be intentional or accidental negligence on the part of the prompt writer to accidentally have their agent write pornography to their website.
The burden of the mistake would have to fall on the person that gave the tool instructions, because it can't know that what it did was wrong. Wrong is subjective in this case. It only did what it did because you, figuratively speaking, encouraged it to.
Yes but I'd actually go even further than that. We've all had experiences where the model straight-up produces gibberish sometimes, right? It's happened to me with a badly configured harness on a local model, and even also on frontier models like when you used to ask them something innocuous about the seahorse emoji.
What I'm arguing is that yes, it's your fault if you misprompt the model and it does something catastrophic. It's also your fault if you prompt it correctly and it does something catastrophic anyway due to some other glitch beyond your control.
The whole discussion started with the point that OpenAI could be financially responsible for damages if their models cause problems for users. If my job is to design a system that works, and instead my system doesn't work, it's my fault regardless of whether I used no LLM, a local LLM, or OpenAI's LLM. Depending on whoever's in charge of doling out consequences, I might get away from it with zero, light, or heavy consequences. But at all levels, I can't reasonably expect to deflect blame onto the model itself.
If future jobs are simply reduced to liability scape goats (or more appropriately reverse centaurs) for management to pin things on then I'm taking up goose farming.
That's more-or-less what you are now, especially if you work at a company like Meta where 1) the guy at the top holds majority control of the company's shares and 2) keeps making massive, expensive mistakes either by accident or design.
This implies that human employees don’t have insurance. But they do. My company’s cyber insurance for example covers breaches due to employee mistakes. Most companies also have umbrella liability policies. It’s just that for now, AI “employees” need a LOT more coverage.
People need to stop redefining and trying to capture the term AGI. None of this is AGI. Not even close. Can it do things an AGI could do? Yeah some of it, but the difference really matters. These things still regularly fail, and gaslight about answers to questions like "how many r's in strawberry" or "s's in espresso".
The last version to fail on those questions was GPT 4.5.
Meanwhile most humans fail to correctly answer how many f's are in the sentence, "Finished files are the result of years of scientific study combined with the experience of many years.".
Nooo. I just asked Claude (Sonnet 5 Medium) how many I's are in assassin, and it said 2. Granted, it got several words correct before this. But no, they still aren't great at this letter-counting thing.
Tokens are the most basic input unit of an LLM. But tokens don't generally correspond to words or letters, rather sub-word sequences. So Strawberry might be broken up into two tokens 'straw' and 'berry'. It has trouble distinguishing features that are "sub-token" like specific letter sequences because it doesn't see letter sequences but just the token as a single atomic unit. 'Straw' and 'r' are two tokens but an LLM is entirely blind to the fact that 'straw' has one 'r' in it.
As an analogy, I might ask you to identify the relative activations of each of the three cone types on your retina as I present some solid color image to your eyes. But of course you can't do this, you simply do not have cognitive access to that information. Individual color experiences are your basic vision tokens.
LLMs see tokens, not words spelled out with letters.
Imagine verbally asking someone who has never seen written text the same question: unless they memorized the answer for the specific word you're asking about, they'd have to guess.
People assume that's the reason because it's intuitive and "strawberry" is one token. But that doesn't explain why those models would also often get it wrong for "StRaWbErRy" or even "s-t-r-a-w-b-e-r-r-y", where the r's are not combined into one token.
We don't know what was going on inside the closed source GPT models, but this paper investigated on some of the open-weight models and found it's not due to tokenization: https://arxiv.org/abs/2604.00778
And circling back around to AGI, tokenization or some other underlying cause should pose no issue. A competent human would think to write a program (ie create a tool) to do the job. It's routine for a carpenter to make a jig.
i asked opus 4.5 what the problem was and it said it pattern matched too much. i asked it how it should do it, it wrote a file that told itself to stop pattern matching. it wrote a file that started with the following and then had an english language procedure for how to count letters. so it knew the algorithm already, but the "instinct" was to pattern match rather than running the algorithm.
CRITICAL: Do Not Skip Steps
Your instinct will be to "just know" the answer. This is how you get it wrong.
You don't see characters. You see tokens. Your "intuition" about character counts is pattern-matching, not counting. It is unreliable.
You MUST execute this procedure step-by-step, writing out each step visibly.
A lot of movies gave the impression that making phone calls and balancing checkbooks would be the easiest tasks for consumer AI to solve, while math and science might require exceedingly advanced AI. Turns out to be the opposite: computers are great at math and bad at conversation.
> "how many r's in strawberry" or "s's in espresso".
And what percentage of your red retina receptors are firing?
The "number of letters" critique was broken before it was introduced the first time. The models were specifically designed with preprocessing to not be able to perceive their input as strings of letters. Blind people are not dumb. (True as a pun and in context.)
Sub-access sensory questions, or do-you-know-a-fact questions (which is what spelling becomes when you can't see the letters, and are not specifically trained to match all token encoded words to their letters) are not intelligence questions.
If the model cannot count letters in a word what happens when it needs to do something akin to counting the letters in a word?
I believe it could easily write a tool to count the letters in a word for frequency, but ... dismissing this as if it doesn't matter seems a bit premature without deeper understanding of bad answers you can get from these tools
But if a model is built to be insensitive to something by design, it isn't a good indicator of its overall capabilities. In fact, it is uniquely poor at testing its capabilities.
I suggest you really try and count the red retina cells that are firing in your field of vision right now.
Or, instead of looking at text, count the e's while someone is talking to you. NOTE: you know how to spell. But try it... Then consider why you can't.
On the other hand, give a text file to a model, and it can count the e's easily. The same information is now in a stable form it can operate on with its higher level functioning.
Who do models and humans have preprocessing layers that strip so much information away? To greatly reduce the cost of operating on information for most purposes - while making other types of operation impossible. When it is presented that way.
If it allows you to distinguish readily between human intelligence and computer intelligence, it seems quite relevant to the question of whether computers have achieved something akin to human intelligence.
It may not be useful for anything else, but at least it can say that.
But it's not relevant as a metric to gauge distance to human intelligence. Humans see individual letters, LLMs do not. If I asked you the relative activation of the cones in your retina as I showed you some solid color image, you couldn't do it. You simply do not have cognitive access to that information. But that says nothing about your intelligence.
A more accurate test would be to give it a list of words (or anything represented as a single token) and ask it how many times that token appeared. I'm sure they have no trouble at that task.
Yes, they have some alien failure modes. But that should be expect, they are an alien intelligence. I might be willing to grant that a lack of ability to reflect on its own level of knowledge is a demerit to it being generally intelligent. But then again it is largely an artifact of training. I suspect if there were a guessing penalty during pretraining they would develop or more readily communicate the strength/reliability of their knowledge.
I am not enthusiastic about criteria for human-like intelligence that imply that dyslexic people don't have human-like intelligence.
[EDITED to add:] I actually don't know whether dyslexic people find it difficult to count letters in words, if they have them already written down by someone else. I suspect they find it harder than people who aren't dyslexic. But perhaps "blind people whose spelling is poor" would have been better; I would not want to deny them human-like intelligence either.
I have slight dyslexia. I can't automatically write double consonants all the time.
But counting certain type of letters is very simple algorithmic task especially in written text. Any reasonably intelligent actually thinking thing should come up with algo and then execute it. Which to me sounds like reasonable minimum bar for general intelligence.
Exactly. It's got nothing to do with sensory input and everything to do with reasoning.
If someone asked me how many f's are in a word I hadn't seen before verbally, then a reasoned response would be that I don't know, but I estimate based on the syllables...or ask them to spell it out.
These are all the sorts of questions where general problem solving works, even if the conclusion is "I don't have enough data to speculate".
So that these models fall apart on it so readily means we're either grossly handicapping then with the requirement to "be helpful" or they just fail to recognize the problem and are just stochastically spitting out a high probability token sequence for the input.
If you aren't an excellent speller and you are asked to count the number of some letter occurring in a passage of text, you will look at its written/printed form and go through the letters one by one. This is, indeed, a pretty easy task and you will probably get it right if you're careful.
The models don't get to see the text written down. By the time your input reaches them at all it's been converted into tokens. By the time they start thinking about it its been converted into embeddings in a sort of concepts-and-word-fragments space.
I do think it's a definite weakness of most LLM systems that they are bad at admitting (maybe because they're bad at knowing) when they don't really know something. (I have the impression that Anthropic's models are better at this than OpenAI's, but that isn't based on careful research or anything.)
What's the actual behaviour of today's LLM systems on these questions when they're allowed to "think"? Someone upthread mentioned that Sonnet 5 at "medium" thinking level mostly gets them right but makes mistakes sometimes. It would be interesting if we could see what its chain-of-thought looks like in these cases.
... I just tried six questions of this kind on Sonnet 5 at "medium" thinking -- this is the default thing you get from free-Claude -- and it got them all right in a way that at least superficially looks as if it's spelling them out and counting. Obviously this isn't enough to guarantee that there isn't anything grossly wrong with its reasoning capabilities in this area, but it doesn't look to me like "falling apart" and it doesn't look like strong evidence that thinking of it as a stochastic parrot is helpful here.
> computers could count the Rs in strawberry since vacuum tubes. that measure is irrelevant.
I don't think it's irrelevant but perhaps not in the way you're assuming. When assessing AGI I'm not evaluating counting characters or even the execution of math operators at any scale or speed. As you observe, computer software from Regex to spreadsheets and Mathematica already handle that well. But AGI isn't about what computers can do, it's about whether AIs can do the specific things which, until now, have been uniquely human capabilities. Like understanding nuanced context and then coming up with novel approaches to solve a new kind of problem not relying on any specific prior training or knowledge (the 'G' is for General).
Most definitions of AGI start from a baseline that already assumes easily passing a Turing test and doing anything via text response that a high school graduate could. I ding LLMs not for failing to count but for failing to intuitively understand the nuanced context of a simple class of problem it hasn't seen in its training data. I fully understand that the reason LLMs fail letter counting is that they operate at the token level. They weren't trained on individual letters first, like human 2nd graders.
The only reason recent LLMs get strawberry and blueberry correct now is that they have those words on their pre-training 'cheat sheet'. However, the underlying fundamental weakness in the way LLM intelligence works which leads to this failure mode still hasn't been addressed. Even when the frontier labs add "recognize any sub-token counting question and write a Python script" to the training cheat sheet so LLMs always pass that test... they'll still be unable to recognize a simple class of problem which isn't on their 'cheat sheet'. As long as that's the case, to me, they aren't AGI because they can't fully replicate human-like recognition of novel problem classes. And it's not just about letter-counting. That gap and others like it lead to many other kinds of non-human brittleness in LLM problem solving. Those are the classes of reasoning, intuition and insight that the ARC-AGI series has been trying to queue up as targets. Not to show how bad LLMs are but to help them be great in all these counter-intuitive edge cases
> planes don't flap wings therefore they cannot fly?
This example still misses my point, which isn't related to usefulness or economic value. I concede that LLMs can have greater utility and economic value than humans on many tasks. The point is most definitions of AGI include something like "can fully replicate all the routine daily tasks done by any competent high-school graduate." That's not related to whether LLMs can solve many high-value problems faster and at larger scale than any human. That was also true of ENIAC in 1946.
The fact an airplane can fly faster and farther than any bird is irrelevant to whether an airplane can "fully replicate all the routine daily tasks done by any competent bird." That's the bird equivalent to most AGI definitions. An airplane can't build a nest or recognize the signals encoded in birdsong.
In the same way airplanes fail the 'bird replacement' requirement, AIs currently fail most AGI requirements only on the terms: "fully", "all" and "any". And in this context, airplanes scoring 15,000% more than birds on 'speed' and 'distance' doesn't matter any more than AIs scoring 15,000% more than humans on 'add 10,000 numbers'. We still aren't near AGI because LLMs cannot fully match any high-schooler's ability to independently conceive new approaches to novel problems not in their prior training data.
which gets us closer to philosophical questions which I'm personally not that interested in.
>In the same way airplanes fail the 'bird replacement' requirement, AIs currently fail most AGI requirements only on the terms: "fully", "all" and "any".
I'm not sure we want a machine that fully succeeds that test.
Planes pass the 'bird replacement' test on the only criteria that matters to us ... flying.
If we wanted nest making planes I think we'd have them by now. Nest making doesn't rate highly on the problems we're looking to solve though.
I don't want a machine that is moody, or depressed or has schizophrenia, which are all pat of the human condition.
We don't need the human "intuition magic dust" to do 99.99999% of useful work.
They're machines designed to do the work we don't want to. That's as "general" as their intelligence needs to be.
I'd prefer if my clothes folding machine did not have an existential crisis.
>I'm confused, is your argument something like "It's too easy so AGI doesn't need to be able to do it"?
Do you possess magnetoreception? a stupid pigeon can "see" the earth's magentic field. why are you blind to it? does a lack of magnetoreception make your intelligence any less "general"
no, you're just blind to it because that's just the way it is.
LLMs are blind to character counting because that's the way they are.
It didn't stop ChatGPT from finding the Jacobian Conjecture counterexample.
Human intelligence and machine intelligence are only going to cross over to a certain degree.
same as plane flight and bird flight are only kinda related.
It matters under the lens of AGI. Artificial intelligence that can meet or exceed human intelligence. Sure, it's good at some things but rather limited at others.
Breadth of capabilities matters... and a promotional video is nice and all, but people are throwing this term around like it's a prize they've won, but they've not gotten there yet.
Is the Turing test focused, perhaps unnecessarily, on deciding how well a computer "thinks like a human"? Should we expand our universe of possibilities to admit that there might be AI that is generally intelligent, but which has some very un-human characteristics?
There never was a clear original definition so people come up with their own ones. But for me the significant thing is being able to do the stuff people do including substituting for them in jobs like inventing better AI. I don't think we're there yet.
The definition that was generally accepted and is often used by places like the FT is intelligence that surpasses human capabilities across nearly all benchmarks. When they first announced Astra they released a statement about how it had solved ten long-standing problems in various fields like mathematics, but it still doesn't really meet the standard definition IMO - and they are gearing up for an IPO - so they have reason to say stuff like that
If you read what ARC-AGI has stated from day one, the tests are designed to stress frontier models in tasks that humans are uniquely good at. When the models get to 100%, the next set of tasks is deployed. When they run out of ideas for how to stress a model (i.e. no more tests), THAT is when AGI is achieved. It's a great definition.
>When we released ARC 3, I got asked, "when do you think a frontier model will saturate it?", and I answered "in about a year, though it depends on how much it gets explicitly targeted"
You're treating an off-hand comment by an ARC 3 researcher as some sort of a precise AI capability acceleration benchmark. Can we leave casual anecdotes (even from researchers) out of the discussions please?
François Chollet wrote in February that he expected ARC-3 to be saturated in "about one year".
"Frontier models today perform very poorly with a minimal harness. However if big labs start directly targeting the benchmark like they did for ARC-2, numbers will go up fast."
Well it's not exactly saturated when OAI refused to use the harness explicitly provided by ARC-AGI. I'm not really familiar enough with the benchmark to declare whether it's a perfect measure for AGI, but I kind of doubt it is.
I think you're misunderstanding. Astra is at the top of the official ARC-AGI leaderboard, with an ARC-AGI approved harness. It's not a harness specialized for ARC-AGI. It just does the same thing the regular ChatGPT interface does: keeps conversation history across turns and compacts when it gets too long. Without the harness, it loses its entire context window every move. That's not how humans work and it's not how any real AI service works.
We are nowhere near AGI. They all talk the same, they can’t help but try to please and affirm us, and if you engage them for too long they become incoherent. They are facsimile machines. They are xeroxing language - but not even, because we can’t even duplicate our results. Too many people mistake the black box quality for magic.
Super useful, incredible tools, but not AGI. Try and roleplay a dialogue with one, make it whatever character and scenario, and see if it can sustain a coherent conversation for more than 30min with you AIM style (aol instant messenger, if that isn’t clear). Expert mode: never correct or adjust it mid conversation.
I’m not even talking about repetition and predictability. It’s nothing like talking to a person. And in a short amount of time it literally can’t form a coherent sentence.
Do you descend into repetitive, incoherent babble after a 30min, incredibly focused conversation? Does any human being without some sort of diagnosis? I assume you regularly have conversations that last more than 30 minutes (work meeting, for instance, which gpt could not participate in as an equal voice by any stretch of the imagination).
It’s an arbitrary number that felt high enough. If I said 10, people would argue that a frontier model can talk longer than that. But I have absolutely watched ChatGPT fall apart that quickly. I’m sure everyone reading this has.
We can nitpick the duration all you want, but we both know it does not take very long for this to occur. It happens particularly fast if you stray from the original topic and/or aren’t using a frontier model.
I work exclusively with OpenAI cloud models, and most high-quality agent harnesses auto-compact and does pretty well sticking to the plan of action. Papers that research the definition of AGI never include context length, and I myself don't see the logic for it either.
I don't think there will ever be universal consensus on....anything. But significant research has been published into its definition, by researchers from the main LLM providers, like https://arxiv.org/abs/2510.18212
Also, I would say SOTA LLM models can sustain a longer, more intelligent conversation than most people. Just in context length alone, they can keep track of more context than humans. But they're nowhere near as efficient the human brain, or the attention to quickly connect and find decades worth of memory like us.
Reasoning is a strong statement here. But it is fair to say that it usually is not intuitive to us.
An example is if I gave you a huge sheet of thin paper (huge so that folding isn’t an issue) - how many times could you fold it in half until you couldn’t physically do it anymore? Could you do at least 10? Try this with random people and you’d be surprised how many say they could do 10 easily.
Or the chess board question. Works to rather get the financial equivalent of starting with a penny and then doubling it for every square on the board or a million dollars for each square? Again, if you ask people to pick one without giving them the time to work it out they will usually pick the million dollar per square.
But you're not reasoning yourself here. In face you're literally parroting what an LLM would do in this situation: you already know the answer so you think your comprehension of this is better than that if others, when in fact it's just simple pattern matching. It's not a measure of intelligence, it's a measure of memory.
I actually started by saying "reasoning" is a strong statement because as you note, it's really not reasoning. It's about intuition. And yes, if you've seen it before then intuition doesn't play a role. But if you haven't had sufficient experience with it, even if you know "double each square", and even if I walk you through the first ten squares on the chess board "OK, now we're at $10.24 cents... and we're at $10 million dollars with the other method -- it's not looking too good for the doubling method is it!?".
But a philosophical question is, with sufficiently large memory -- does almost everything just reduce to a measure of memory?
If something at rest is accelerating at 9.8 m/s^2, how long in seconds will it take to reach 10% of c? Answer to the nearest order of magnitude - will it take approximately 1000, 10k, 100k, 1000k seconds?
I’m sure you know this is an exponential growth question but have no intuition of the answer.
Knowing exponents and how to apply it is not the same as having any intuition about what the actual value of a certain exponential function will be at a certain point and when it crosses a threshold.
> I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.
To me AGI is all about the "G" general (we already had the AI part). General meaning universal, everything. It's not a function of knowledge or specific hardcoded tests, it's that you could give it a test it's never heard of before and never been trained on and it would ace it (it might need a lot of time).
Currently LLMs can't even really learn within a conversation, they can add a note to context and try to not drop it. Example things an AI cannot do yet (but maybe someday will):
- write a well-received book, write a best-seller
- come up with a new company idea, Run that company
- actually have a decent conversation, maybe someday talk somebody out of suicide effectively
- come up with its own ideas or theories that nobody else has presented
- understand the stock market well enough to trade better than an index fund
- be an expert Game Master in a TTRPG (making no mistakes, getting a read on the players' fantasies, calibrating difficulty in response to emotions)
- come up with a theory of what makes games fun, make a popular game
- be able to sort through research and come to conclusions on complex geopolitical/sociological topics (e.g. theorize on whether AGI will result in mass poverty or mass abundance and be able to argue persuasively)
- be able to articulate what it knows, what it doesn't know, and what information it would need to have to answer complex queries
- exhibit metacognition (thinking about its own thinking) and self-optimization
- wonder about things
- observe contradictions and ironies in the social-consciousness, do a standup routine that makes you rethink how you look at things
However most humans can do at least some of the things given they spend the required effort.
Some problems presented needs a very large context and some are not much solvable (e.g. trading) since market responds to traders' actions, as well, making it effectively an oracle problem (of computation).
On the other hand, we must be aware that these models are static, and they indeed stop when nobody asks something or requests an action from them. However, brains in nature never stops. Wonder, daydream, sleep, self-evolve, clean up and eliminate memories and views and much more.
But this is assuming the model is the entire story. The original comment you were replying to pointed out that the harness is just as important.
The hardware of human intelligence is not a singular thing that is uniform throughout. You cannot take the prefrontal cortex white matter out of someone's head and say you are holding a person. Much of the parts of our brains that enable much of our intelligence, is made of different specialized stuff. The visual cortex and sensorimotor regions aren't only there for input and output, they are used by the more thinky parts of the brain to do visualization and spatial reasoning. The cerebellum contains billions of neurons making little oscillator circuits and PID-like self-regulation machines that help make muscles do what they're supposed to, but also provide attention and time perception.
Heck, our brains contain language models, that train themselves up based on a glut of data over a span of about 10 years, and then they become more or less set in stone for the rest of our lives. Of course we can learn languages, but the "Critical Period" is a very real thing that produces a permanent architecture for some grammatical structures, or things like the ability to partition a lexicon by gender for faster lexical access which cannot be learned as an adult if your native language did not have gender.
I'm not trying to make a direct analogy, the point is that the language model doesn't need to be fully "generally intelligent" all on its own for there to exist a general intelligence, because the language model can be part of a generally intelligent system, which can do things like form, recall, and manage memories which are by now a standard feature in basically every chatbot.
That's true, but my general gripe is the static nature of the whole system. Even if humans' neuroplasticity reduces with age, it never goes to zero. I started to learn English at a very early age, yet my ability to communicate with it soared around 20, because I started to use it more and more.
Same with instruments. I started to play instruments at an early age, but started to play guitar around that age. Well, I'm not a virtuoso, but can play and more importantly can improve.
These AI systems we built are static things. We generally try to make them more intelligent by augmenting the context they can see, but the model doesn't evolve in every turn, for example.
Intelligence is a multi-faceted and multi-input construct, that's true, and GPT-6 may do amazing things w.r.t. other models, I didn't try it yet. OTOH, my main call is to remember that these are still static algorithms fed with enormous amount of data. They are more rooted on statistics rather than fixed inputs. In short, they are still fitting to the frame of "advanced search".
In general, I'm not against the tech, but the hype. I have other gripes about how AI is being built, but that's not subject of this comment.
> On the other hand, we must be aware that these models are static, and they indeed stop when nobody asks something or requests an action from them
The parent commenter noted:
"if a harness alone can help the LLM saturate the benchmark with a near perfect score then the combination of the two must still be effectively AGI"
Harnesses absolutely can enable models to continue thinking about things. And LLMs do wonder and explore weird ideas like daydreams when you allow them to do this.
However, this doesn't change the fact that you are pumping more and more tokens to a static model's context window, even if you do compaction, the model is not more intelligent than previous turn.
Most intelligence researchers would agree that people seem to have a genetic cap on their intelligence. While someone can underperform their intellectual potential with an upbringing that doesn't adequately enrich their minds, it's near-impossible for humans to become more intelligent through reading, studying, etc.
When humans learn we gain knowledge, not intelligence.
I think the only real difference is that we humans are born lacking a lot of initial knowledge/data which means we have to go through a decade or more of education to reach our potential intelligence. LLMs on the other hand come pre-loaded with that knowledge.
Passed this point, wherever knowledge is passed in as context or stored in the neural net I don't think is that significant personally. I'm of course not suggesting we're exactly the same as LLMs and there is no noteable difference, I just don't think continual learning is as important as some suggest it is – at least assuming a model is deployed with adequate training such that it reaches its potential given it's size + architecture.
I'm guessing their defn of AGI is something like the sum total of all humans' abilities? Still though some of those tasks (e.g. beat an index fund) may very well be impossible, and worse yet a lot of those tasks are not coherently defined.
The point is replacing humans, and for that it only has to equal them at lower cost (including factors like not needing sleep and being easily clonable). The outperformance lies in the cost savings, not necessarily in the intelligence.
To me all this makes the label of AGI completely meaningless.
What AGI has always meant (eg. in 2019) is Artifical General Intelligence.
Artificial -- something made by humans instead of occurring naturally
General -- not confined by specialization or careful limitation
Intelligence -- the capacity to learn, reason, solve problems, think abstractly, and adapt to new situations
Basically, the metric was that any healthy adult human on the planet represents a general intelligence. This has certainly long been reached.
Also some of the stuff you're listing has long been solved as well, such as listing what it knows and what it doesn't know, and what information it would need. Other is just poorly defined: "be able to argue persuasively". AI can certainly write an argument on almost any topic that would pass any University homework in 2019.
Your examples are things that most humans cannot do, or things that AI can already do. For example most humans, even most intelligent humans, could not write a well-received book, run a successful company, or make a popular game. On the other hand, AI can absolutely sort through research, draw conclusions on complex topics, and argue them persuasively. Likewise, I don't know what you mean by a "decent" conversation, but millions of people converse with chatbots daily, so I don't know why you say AI fails to meet that bar.
I fail to see how what you describe is any better than the old autocomplete-on-steroids comparison. Could the human mind learn to spell every word properly? Yes. Do most (or any) do it? No. Does that mean a human can't?
I think what OP was drawing a comparison to is that AI right now could not come up with an award winning novel from the spark of some creative notion and working up from there, as opposed to just mashing together what has already been done and calling it a day.
I don't think that's strictly true, as I can give it a new gui or tui program it wasn't trained on and it will learn it. Unless you're talking about general abilities like sight, but the same is somewhat true of humans.
If you consider the data on which an LLM was trained on to be points on a very highly multidimensional object, the claim is that the LLM can interpolate a convex hull spanned by those points, therefore recovering a subset of consequences attainable from those points. Obviously this hull includes completely novel points that were not present in the initial data set, so the output of the LLM goes beyond its initial training. And yet, there are clearly points outside a convex hull spanned by any finite number of points, such that we can imagine not all possible outputs are attainable using this method.
The claim is furthermore that truly original thinking, the infamous leaps in understanding and creativity, happen by attaining points outside such a convex hull.
It's hard to rigorously verify or disprove this claim. Hopefully this helps build an intuition of why the claim is not as shallow and obviously wrong as it may seem initially.
This makes me realize there is a higher bar we need to achieve with AI still. The ability for the model to evolve through interactions more on a hourly or daily basis. The models are accelerating but inference doesn’t modify the model.
But you have to acknowledge how uneven the playing field is. The AI has read every book that's ever been written, and can spend hours of compute time in a few seconds.
I think if a person had those same advantages (e.g. could spend 5 hours thinking about what to say next) we could all hold outstanding conversations, or if we had read every book ever written I think many of us could write a very popular book, if we could read every singe company's P&L statement in a few seconds we could invest better than an index fund.
What I'm pointing out here is that these models appear to be intelligent when they really are simply unimagineably knowledgeable. When you drop the time-constraints it starts to become more and more apparent that human intelligence scales better with time than AI does (much in the same way AI can burp out tons of code but make your codebase entirely illegible within a matter of months).
Perhaps to simplify: my notion of intelligence is how much can you deduce with a constant set of starting context
> I think if a person had those same advantages (e.g. could spend 5 hours thinking about what to say next) we could all hold outstanding conversations, or if we had read every book ever written I think many of us could write a very popular book, if we could read every singe company's P&L statement in a few seconds we could invest better than an index fund
I’m not so sure of that - to get average outcomes in these fields it’s a matter of time, to get above average or extraordinary, you need talent/intelligence/taste.
So now the bar is not only to be at the level of a human, but to achieve it without the natural advantages of being a machine, while still having the disadvantages (lack of embodiment, etc.).
If you went back to when I started working on AI stuff 20+ years ago and described the capabilities of GPT-3 to people, the overwhelming majority would say it's a form of AGI. It's incredible how fast and far the bar moves.
I would be willing to bet that any human for which we spend $100billion - $3 trillion (depending if you want to count single corporations or global totals) on in an attempt to make them as capable as possible would be able to reach all of those levels.
> A quick search suggests that the most expensive education in the world is something like $100k.
I have a kid in an American university right now, and a quick search of my bank account statements confirms that there are far more expensive educations in the world.
But that list is extremely ambitious. Write a best seller, make a popular game, come up with a truly novel theory, consistently outtrade index funds.
That's top 0.001% human stuff, I don't think you can take just any person and get there through education alone, it takes extreme talent and dedication. There's also diminishing returns when spending on education, it doesn't just improve linearly.
The marginal return on education spending decreases fairly quickly, but obviously becomes zero at the point by which there are not enough hours in the day/year/decade to cover every single topic that humans know about - no matter the talent or resources available to the student.
except you are missing the one versus many argument here. sure we could make one human much smarter, could we make endless copies with the same intelligence? no
For $100 billion we could pay ivy-league level tuition for a million people. You don’t think investing that much in education would yield some good research or companies?
You'd get rapidly diminishing to zero returns after the cost of university a few times over. Every dollar past that would produce no performance gain beyond that.
Are you imagining artificial augmentation somehow? Purely through tutors or training programs we seem pretty limited. Otherwise billionaires (or even multimillionaires) could have far more consistently successful kids.
Don't the children of the wealthy famously have a tendency to be successful? Or have I badly misinterpreted the last several thousand years of human history.
To really drill down into that I would think you would need to figure out how many millionair children get tutored vs how many get spoiled.
>be able to sort through research and come to conclusions on complex geopolitical/sociological topics (e.g. theorize on whether AGI will result in mass poverty or mass abundance and be able to argue persuasively)
i'm not sure what makes you think AI cannot do this already. in my experience, this sort of deep research is something AI is quite good at.
I understand that what it came up with sounds impressive (especially since I know 0 about Myanmar), but on the topics I do know about its analysis routinely have very fundamental problems (even this Myanmar analysis has % that add up to > 100). There's a chance it's just parroting the majority opinion on Myanmar, or making stuff up (and perhaps you could ask it to write a strongly worded opinion in the other direction that would sound equally plausible).
For example I asked it to do a full analysis on the AI bubble, and a full analysis on the risks of Glyphosate, and it came up with a lot of things that sounded credible, but within a few minutes of questing was admitting it hadn't even really checked for internal consistency in its positions, and even doing a 180. It certainly was much faster at gathering sources and reading but it fundamentally doesn't seem very effective at creating a consistent worldview.
And of course the funny thing is it says it did a 180 on one of these topics, great, except whatever it concluded will be discarded because it cannot learn. It's just bonkers to me pretend this is AGI, it probably couldn't even hold its own in this very discussion.
Regarding this specific case, the analysis is sound. It is the majority opinion on Myanmar, but that majority opinion is held for good reason, e.g. China's desire to keep Myanmar together, complicated ethnic boundaries, desires of current EAOs.
The % responses adding up to over 100 makes sense, because the 25% outcome (formal secession attempt) and the 10% outcome (completed secession) could both occur, so it's implying that if secession is formally attempted, there's a 40% chance it will be completed. The numbers are broken down more clearly at the end - it is a bit confusing at first glance though.
> You want a computer program to be able to take a single phrase and execute decade long journies?
In my opinion that is exactly the point missing from AGI: the fact that you still need to prompt it. As long as you have to ask for something, is not general.
Autonomy is not the same as general intelligence. We already have all kinds of fully autonomous technologies that are nowhere close to generally intelligent. Plus, at a certain level of abstraction, human beings also need to be “prompted” to some extent by stimuli. And this is the funny thing about general intelligence as a concept: most of the definitions that come close to internal coherence rely on references to human intelligence, a concept we feel like we understand because we all live it all the time, but whose actual nature and structure is extremely slippery.
I don't get it, human employees frequently need to ask for directions too?
They often act on their own, too, and get things wrong a lot. The reason it works is because of all the systems of laws and institutions we have built around humans, not so much because human minds are special.
It sounds like what you're saying is that AGI should have some sort of free will. I'm not sure why you would add that as a requirement. Could you expand?
I think they merely want something with a functioning long term memory. Something that can exhibit growth past the first 5 to 10 human-equivalent hours working on something.
Current LLMs are worse than most dementia cases, reaching "peak domain skill" pretty much immediately.
>You want a computer program to be able to take a single phrase and execute decade long journies?
> Who will be responsible for the outputs and side effects of such a closed loop system?
Itself. That's the point. We can do it. Until it can met that bar, it ain't AGI. That's always been the bar.
"Being a person in all of its aspects" isn't the same as "generally intelligent". The latter is at best subset of the former, and it's also easy to imagine a system that is more generally intelligent than humans, without being a person. See also discussions of the personhood of various animals who are less intelligent than average humans.
- Amazon is full of AI books, and they're clearly making money. AI has won multiple literary and artist awards.
- Okay, it's not a "new company" idea, but VendingBench is all about ability to run a company
- Plenty of people disagree with you on conversational quality; see "AI Boyfriends" etc.. (and it's not hard to find people who consider it uniquely valuable for discussing mental health)
- "come up with its own ideas or theories that nobody else has presented" C'mon, seriously? Solving a half-dozen hard open math problems wasn't enough there? What the heck counts as "it's own ideas or theories" at this point?
- plenty of evidence that custom models are starting to do well on the stock market, although I'll admit we're a year or so from any solid proof, since you need a track record to really make the claim
- LLMs have been capable of being a GM for a TTRPG for over a year (although like humans, they make mistakes)
- Okay, conceded, but humans tend to take years and large teams to make a game. Even if the capability existed today, it would take a while to actually build, test, market, etc.. - all made much more complicated by gamers being largely opposed to AI art styles, etc..
- "be able to sort through research and come to conclusions on complex geopolitical/sociological topics" - uh... did you mean to say something else, because "come to conclusions" is... like, LLM 101?
- Hahaha, have you met humans? We definitely cannot do that.
- Uh... thinking about it's own thinking is trivial. Most LLMs these days are built using LLMs, so uh, self-optimization seems nailed, too? We just don't let them do it unsupervised.
- LLMs fucking love to wonder about things
- "observe contradictions and ironies in the social-consciousness" really seriously have you actually used an LLM recently? I think you would find it remarkably enlightening.
I would bet that llms have talked plenty of people both into and out of suicide at this point.
That nitpick aside, I think that's an excellent list. Especially being able to articulate what it does and doesn't know, or how confident it is. That's something that naively sounds pretty simple, but clearly isn't. And it's something humans aren't great at either (see: Dunning-Kruger), but so far LLMs don't even really have the capability to attempt it.
Ordinary people do these things all the time. There are new companies made every day, new books top the charts every week/month/year, same for music.
People have decent conversations every day. Ordinary people sometimes do have to talk someone out of suicide.
Yes, average humans are not beating the stock market. But the average human is a bit better than you give credit to.
Please consider the context of the question. An artificial intelligence only needs to have the cognitive abilities of a random average human in order to be "AGI".
The average human has never published a bestselling book. A person who has published a bestselling book is an above-average writer. And, therefore, an artificial intelligence capable of writing a bestselling book would be above an average human at the task of writing books. Therefore, somewhere beyond an AGI.
Attempting to redefine AGI to "being better than most humans at most tasks" is moving the goalposts towards artificial superintelligence.
Are you mistaking "average" for the literal average number of people? I am referring to average as in ability. Most authors are pretty average people, sure, the above average ones write war and peace etc. But most books these days are not from these extraordinary people.
Indeed. And there is a wide gap between "being as cognitively capable as the average human" and "doing things that no human can", as you put it. So it is important to be realistic about what things the average human is capable of doing, and writing bestselling books is not one of them. Just as an example.
And in order to describe the capabilities of what is AGI and ASI, it is useful to look at examples of what humans can do at a range of capabilities, with some nuance. It is not black and white.
It costs money to train each single human, who is then only productive for a number of years until age takes its toll. Once you have trained one software system, the marginal cost of producing a copy approaches zero. Every subsequent improvement can be broadcasted in a matter of seconds across thousands of data centers. In addition, software does not get sick, age, or die.
So the goalposts have moved to include continual learning.
In a sense I think no one will agree on a definition of AGI until it becomes impossible to construct any benchmark under which an AI underperforms "average" humans. That or it's defined retrospectively, after it's overwhelmingly obvious it met any such definition.
I hardly think it’s fair to label an objection so old that Turing included it (and discussed it at length) in the list of objections to thinking machines in 1950 “moving the goalposts.”
> These arguments take the form, “I grant you that you can make machines do all the things you have mentioned but you will never be able to make one to do X”. Numerous features X are suggested in this connexion. I offer a selection:
> Be kind, resourceful, beautiful, friendly (p. 448), have initiative, have a sense of humour, tell right from wrong, make mistakes (p. 448), fall in love, enjoy strawberries and cream (p. 448), make some one fall in love with it, learn from experience (pp. 456 f.), use words properly, be the subject of its own thought (p. 449), have as much diversity of behaviour as a man, do something really new (p. 450). (Some of these disabilities are given special consideration as indicated by the page numbers.)
Just because a condition is new to you doesn’t imply moving the goalpost. People have been putting forward continual learning and similar conditions like autonomy since 1950s.
Only someone who doesn't do any work of any meaningful difficulty could think these models have anything to do with AGI.
Today I spent half a day trying to solve a moderately interesting software engineering problem. I was switching between GPT-5.6 Sol and Fable 5.1 to check each other's work in Cursor.
And the result was gradually driving me insane. As the models struggled to find a solution that would actually work, they dug themselves deeper into a hole. The work grew in complexity beyond my ability to understand what's happening and recover.
At some point, when I felt like throwing the keyboard out the window, I just gave up. Tomorrow I'm starting from scratch, having burned god knows how many tokens and hours of my life.
But sure, they can create a decent website or CRUD app, so they must be really smart.
But that happens with humans as well. You are having the same experience with an AI that many managers have with their direct reports.
The smarter AI gets, the easier it becomes to move the AGI goalposts. Seems at this point there are people who will refuse to call anything less than omniintelligence AGI.
(And then the excuse will be, but it’s not omniscient! And even if it were, is it omnipotent?)
My success rate for solving software engineering challenges encountered in my day jobs has been near 100% for my entire career. I can only think of a few tasks I kicked back and said they were impossible. For example, after trying to get a signal processing system working reliably I decided to sit down and calculate the actual limits of the channel we were sending the data over and found that from a basic estimation it would not be possible to do. In start ups you don't really get to get stuck in a spiral and not fix things.
I find agents often get into these cases during research tasks.
Yep, people are typing comments with a computer that is powered by several layers of software that will be stored on another computer powered by several layers of software to be read on a computer also powered by layer of software. And then they hope to make the argument that humans cannot produce software.
yeah for me it's "I wanna add this new thing to an existing system" and the AI responds "we should just add some arbitrary state here to facilitate this feature".
The real issue is the existing system needs to change entirely to facilitate, I know this, Good developers know this, The AI however knows the shitty solution would solve the immediate problem because it's been trained on shitty solutions.
the problem could simply be the AI doesn't have all nebulous loose context I have about the goals of the project and future plans, but I would have to write a novel to give it that context.
> The real issue is the existing system needs to change entirely to facilitate, I know this, Good developers know this
1. I’ll often include boilerplate in a prompt to tell it to make the broader fix. [1]
2. However, a top HN AGENTS.md post 11 days ago included the standard guidance “As much as possible try to minimize the number of changed lines when implementing a feature.” I.e. some devs want LLMs to avoid broader changes and so some of that likely makes it into the training, even if others like us want the opposite.
[1] As far as whether my boilerplate is effective, I don’t know.
The default case when implementing a change should obviously be to make changes with as minimal a blast radius as is reasonable. This is basic software engineering and current LLMs fail it. LLMs also don't have a good idea of what is "reasonable" and a human judgement call is needed.
The case where a (sub)system needs a complete rewrite to admit a feature without incurring too much technical debt should be the exception. When exactly to make that exception is something that clearly currently requires a human judgement call, as models aren't yet nearly smart enough to make such calls.
Humans dig ourselves into holes as well. Sometimes more intelligent humans are better at realising they are digging a hole and clamber out, but sometimes they just dig deeper.
And: Is your work more difficult than finding proofs of or counterexamples to decades-old open problems in mathematics?
I was using an AI to help me set up a container to be used as the Nix build environment for another AI. This build environment would not have Internet access. I was having it base its approach off a previous container used for a Stack build environment.
In its initial analysis of my proposed strategy, it insists as its premier point /against/ the strategy, "you will have to rebuild the container every time your flake.nix changes."
Two head-slapping errors of judgement in saying something like that:
(1) The Stack solution is identical. Change stack.yaml, the container must rebuild.
(2) It is not physically possible to do better than this while insisting on an internet-free environment.
So on this point, it was just parroting advice irrelevant to the context at hand. LLMs always have such a bizarre mix of technical knowledge and lack of good judgment.
Not the person who asked, but yes, that sounds like a small lack of judgement, but hardly a huge problem. You just correct it and move on.
I work on some reasonably sophisticated stuff (not inventing a new form of compression sophisticated, but still) and I just don't seem to encounter so many of the issues people talk about with these models.
It's really hard to say why, everyone uses them differently. I generally start complex tasks with a "here is what I am trying to achieve as a high level, here is a file that lists the technical constraints, here are my initial thoughts, here is where I am uncertain, what am I missing, lets have a deep back and forth discussion about it with the aim of ...."
Not always the same prompt, definitely not when the task is simpler, but this usually gets me to a good place before it I get it to write any code.
And our repo has very strong opinions and guidelines for how testing is done - we're lucky enough that we write mostly single threaded low latency code, so testing it end to end is very easy.
Wouldn't "general intelligence" require so much more than scoring well (or even amazingly) on benchmarks?
Like what about having some "AGI model" embodied in something (maybe humanoid), and test it by having it step in an assortment of cars and park them. Does bodily-kinesthetic intelligence account for nothing? Humans are intelligent creatures and can dynamically adapt to the physical shape of a variety of vehicles and their movement characteristics. And there's so many things like this that are extremely basic, which some people dismiss since practically every human has the capability to do it, but actually requires a high degree of intelligence.
What you've described is just a new benchmark, though. It'll be called CarParkBench, various embodied LLMs will then be run against that benchmark, and some will score better than others.
I do see where you're going, but that's already what's happening: we have so many different benchmarks because there's no real single way to test for general intelligence.
Also, it takes a human probably at least a decade of world experience, growth, learning, etc, to pass your benchmark. I'm quite confident that it will be very soon that an embodied LLM will pass your new benchmark, much sooner than a human would take if born today.
I think the issue is less about creating a new benchmark and more that the existing benchmarks shouldn't be called anything related to AGI unless they measure AGI.
If a model couldn't go to work as e.g. a first year apprentice plumber on their first day and perform anywhere remotely close to the median, but can pass a benchmark that claims to measure AGI, the benchmark is wrong and the model is not exhibiting general intelligence yet. ApprenticePlumberBench sounds like it's genuinely better than ARC-AGI at measuring AGI and that's a bit silly.
(Edit: I wrote ARC-GIS the first time around, for some silly reason)
It's not measuring AGI at all, it starts from human "core knowledge" so it is parochial. It is made of tests that still fail so by definition next version will also start low. Moving goalpost.
This is a big reason why I feel like even though LLMs are _effectively_ AGI in some regard, they also are a hack around what most people figured AGI would look like before the advent of LLMs. Humans can do metacognition, output multimodally at the same time (verbal _and_ physical intelligence go together to produce an expressive face while one talks), have a good sense for what they do and don't know, continuously take in and respond to the world around them in a (mostly) uninterrupted fashion without "turns", learn knew knowledge and retain it for their whole lives, etc. When you reduce a human to a text generator, yes obviously SOTA LLMs perform way better, but rather than invent something that can operate as an always-running "being", we've grafted a harness around an intelligence that is bound purely to speak only when spoken to. Maybe organic intelligence is already that, playing out at a super high refresh rate, but I don't know.
Current AI is arguably much more capable of multimodal output than humans. It can produce an incredibly vast variety of audio, images, and video. Humans are limited to producing the sounds we can make with meatflaps in our throats, and contorting various parts of our bodies to produce crude symbols and shapes.
(Very capable!) Embodiment, persistent operation and continuous learning are indeed things that still set us apart from AI. None of those are fundamentally difficult to solve, though.
More importantly, none of those are particularly relevant for being "intelligent": If a criminal threatened to kill your family unless you solve some difficult problem that requires only intelligence and you could choose any single person, animal, or AI to help you with it, which would you choose? Be honest.
The thing I trust the most to solve tricky problems reliably is a specific very skilled programmer I've known for twenty years.
I wouldn't say he never makes mistakes, but his success rate is a damn sight better than any LLM I've ever interacted with (and I drive Opus daily, due to corporate demands to use LLMs).
Same partial answer as I gave your sibling commenter:
"OK, so the task ends up being to write a 40-page analysis of the result of a specific experiment in quantum chromodynamics, to be finished within 2 minutes. Does your choice for a human work, or would you have rather chosen any of the frontier AI assistants? Be very, very honest. Remember that the lives of your family are on the line and nobody will judge you for a lack of allegiance to humans."
This task is intentionally designed to ensure a human cannot do it.
The initial scenario is utterly, insanely absurd to begin with, but I tried to go along in good faith and gave you the true answer.
The result was a bad-faith rhetorical trap, so I'm done with this thread.
In another attempt at good faith, as part of bowing out I will add some actual response to your anti-useful cheap rhetorical trap:
I do not trust LLMs to get things right in high-stakes scenarios. I have seen the current models spit out falsehoods and errors regularly in the handful of fields I have expertise in, and have no reason to think they would do otherwise outside my expertise.
The scenario you describe is an absurd fiction, and no human making the absurd threat could evaluate the paper in less than hours (realistically even an expert would need days, and a nonexpert could not do it at all [short of becoming an expert]).
So, there's no point trusting a bullshit machine to save my family - it might very well get them killed, and whether it was right or not, what would actually matter would not be its correctness, but what the presumable bullshit machine evaluating my offered input spits out.
So, the best move I could realistically make would be to put a stab at prompt injection into the input.
For that job, I probably would actually prefer aforementioned programmer over any other option, come to think of it - I suspect he'd have better success than even another model (especially considering the safeguards the models no doubt have to try to keep users from using the models to inject other models).
Again - I'm disappointed in your worthless rhetorical cheap shot.
I suspect you'll have much better success convincing people LLMs are intelligent if you engage in good faith, listen to their perspective, and address their actual thoughts, instead of devising the sort of inanity that comes out of high school debate clubs, where people literally want to score points instead of find truth.
> This task is intentionally designed to ensure a human cannot do it.
It is one of a vast array of things the hypothetical kidnapper could come up with, some of which humans I agree will do better at (currently) and some of which AI will do better at. We clearly agree that in that array there is at least one task that a frontier AI would be better at than any human you could pick.
It is a thought experiment, so there is nothing fundamentally wrong with it being extreme or unrealistic (thought experiments very often are), but for the sake of goodwill let's 'weaken' it a bit: the AI or human always has an hour to come up with the answer, the question and answer are in their preferred language, and the answer fits on four pages. You can't help them, though; The criminal 'prompts' them. They can use the internet as an informational resource, but they can't communicate/ask for help/post anything (with the spirit of this being: no loophole in letting somebody else do the task or parts of it for them). And of course all subject matter of all complexity is fair game (including but not limited to quantum chromodynamics experiments).
Given that situation, do you think the programmer you mentioned would be more successful than a frontier AI in more than 50% of the possible intelligence tasks?
Edit, addendum: Please, if you can, also let said programmer read this thread and give his opinion on it. It sounds like he would have interesting things to say on this.
Before "AI," humans have created a vast array of "multimodal output" (computer art, instruments, dance, architecture, etc.). Why are you giving the AI a harness and a plethora of tools and not the human in this comparison? Without these, the LLM too would be utterly useless.
Also, that situation is extremely contrived. If a criminal threatened to kill me if I misspelled a word, I would choose a dictionary. If a criminal challenged me to predict a next token, I'd choose the LLM. For all real precarious dangerous situations, I would obviously choose a human. Like immagine the hilarity (or tragedy) that would pursuit if ChatGPT tried to handle a hostage situation or a plane hijacking.
and those meatflaps are normally called vocal folds/cords btw
> Why are you giving the AI a harness and a plethora of tools and not the human in this comparison?
I am not. Multimodal models generate that output directly, without tools. Which 'tools' does AI use to generate all those images, songs, and videos do you think?
> Also, that situation is extremely contrived. If a criminal threatened to kill me if I misspelled a word, I would choose a dictionary.
Of course it is contrived, it is a thought experiment. Does not make it less valid. It is essential that you don't know what task it is going to be, just that it is a task requiring a lot of intelligence. This way question dodging loopholes like "I'd choose a dictionary" are impossible (and people will always try to find some cheesy exit rather than facing reality). You have to commit to something or somebody that has broad and general intelligence; you do not have the luxury of choosing the perfect tool for a very narrow task.
> For all real precarious dangerous situations, I would obviously choose a human.
OK, so the task ends up being to write a 40-page analysis of the result of a specific experiment in quantum chromodynamics, to be finished within 2 minutes. Does your choice for a human work, or would you have rather chosen any of the frontier AI assistants? Be very, very honest. Remember that the lives of your family are on the line and nobody will judge you for a lack of allegiance to humans.
Again, don't go for shitty loopholes. Engage with the thought experiment in good faith and thus as it is stated, not some conveniently distorted version of it.
> and those meatflaps are normally called vocal folds/cords btw
What? Next you're going to tell me that meatspinner is also not the name for the human male reproductive organ.. Maybe I need to get a refund on my Temu Gray's Anatomy.
That said, a look at the state of self driving and the recent robot olympics shows that advancement on that has accelerated enormously, though whether it's reflected in any of the LLMs is something else entirely.
>AGI has a pretty precise definition, covering only cognitive tasks.
OpenAI's own charter defines AGI as "Highly autonomous systems that outperform humans at most economically valuable work". This is actually fairly sensible and involves obviously a ton of non-cognitive, emotional, social and physical activity. In other words, if you can replace most or all human beings with a machine, you have something that's generally intelligent.
That's obviously not even remotely where we're at, AI chatbots do well on narrow usually text based or programmatic problems, but can't even replace a barista or a plumber.
There is rapid progress in generality in humanoid robotics though. I think within the next year or less we will get the ChatGPT moment for humanoid robots. If you look at progression of capabilities such as the recent Skild AI demos.
There are simply no technologies today that can replicate the density, precision, and versatility of human touch sensors. Until then, there is simply no way to create generally capable robots that can operate at the level of a human.
And unlike LLMs, advancement is held back by physical limitations like materials science, so progress has been and will continue to be much slower.
What task do you think that humanoid robots can't do? Also, we don't need fully equivalent touch to get useful performance.
If you look you will see a really broad range of tasks accomplished already, including thing like manipulating screws, picking up pills, inserting wire harnesses, folding clothes, putting away dishes. And there are several companies with built in or component advanced touch sensors like Figure or leading edge touch sensor companies like SynTouch and GelSight.
Peel an orange? Crack an egg? Thread a needle? (Heh, drive a car...) There's a huge range of tasks that a non-specialized, general purpose robot simply cannot do yet. I'd be easier to enumerate the things they can do than the things they can't given the current state of the art.
Sure, build an orange peeling machine and it'll do great. But that's not what we're talking about here.
As for those demos videos we often see, those are very highly choreographed demonstrations. Show me a real life humanoid robot operating free form on a factor floor and doing those things and I'll be impressed.
And to be clear, this is not meant to understate what's been accomplished. I'm just saying the path for advancement is a lot harder and based in physical rather than computational limitations, which are much harder to overcome and go much slower. We simply cannot look at the growth curve of LLMs and expect robotics to advance at the same rate.
peeling an orange and cracking an egg already demonstrated. Figure 02 worked on BMW's actual Spartanburg production line for about 1,250 hours running 10-hour shifts.
keep paying attention, you will see how wrong you are about it being physical limitations as the physical AI continues to rapidly improve.
Surely part of the problem is that intelligence seems to be implicitly conditioned on embodiment, to the point that "covering only cognitive tasks" seems inherently ill-defined or arbitrary. Everything we do is a cognitive task. At some point, the criticism will be "sure, it can solve research-grade math problems, but it can't fold my laundry".
Even our large language models have an implicit embodiment in the domain of text (and more recently, multimodal inputs). That seems sufficient for certain things, and insufficient for others. I suspect that AGI that does everything a human can do eventually turns out to be fairly analogous to humans in terms of sensory input and domain output, even if the scale is radically different (e.g. thousands of robots uploading (touch, sight, audio, smell, etc.) sensory data to a single model, and each being actuated individually).
Small comment regarding the ARC-AGI-3 scorecard: the ARC folks published a blog post as well [1], reporting that without the custom harness, Astra (max) achieved 62.7%, which is still a huge jump from Opus 5, albeit not at the 99.9% that OpenAI self-reports with their harness.
If you define AGI as "can do the work of a human sitting at a computer, end to end", then I'd say comparing yourself to it on a specific skill is the wrong test. Can you hand it a role and walk away for a day/week/month?
I can’t yet. I think that I'd want at least two things it doesn't have: the ability to retain what it learned yesterday (without me carrying it in the context window and thus micromanaging it), and the ability to prioritize correctly, i.e. tell which of the n things it could do next is the one that actually matters.
Can't say for sure that those are enough, but not having them seems to be most of why I still have to "babysit" these incredible tools.
I think fundamentally it is that. The ability to retain information.
Like given a specific task it can do a thing amazingly well, but can it recall a thing. Its memory seems like a giant filing cabinet and it has to go scan like 20 million tokens worth of memory to recover things previously talked about.
Human memory is more graph like, we don’t recall things exactly, but one thing links to another, we create a pattern of a thing, we mark what is important, and overtime what was important degrades or becomes less so.
I feel like what makes it lack intelligence is it never seems to learn. Like it kind of does, but then doesn’t persist once too many other things are learned.
I’m sure they’re probably working on this, but I feel like that is what I want far more than even better models, is a better memory system to recall and forget things that the models work on.
To be fair, I think you can't hand a role over to someone you just hired and walk away for a week. No matter how much of a SME they are. Let's not forget human onboarding takes months. With the advantage of their knowledge not going into the void multiple times a day. That's likely the last missing piece, a solid system of memories that produces the same effect as short/long term memories in a person.
I agree, we're not at that kind of long horizon capability yet. You still need a human in the loop to do manual testing. For whatever reason, we managed to automate the skill before we managed to automate the focus.
So, for now, humans need to stay in the loop and do low skilled labor to keep the skilled work the models do from going off the rails.
To each their own. Personally I will start feeling the AGI as soon as we move from chatting about benchmark results to learn that some lab just announced the discovery of tens of novel treatments for rare diseases.
Maybe I'm too boring but it seems quite pointless to have this same prediction game every time a new model is released.
Whether AI is AGI does not depend on the speed at which it operates/thinks. Clearly all the theoretical work done by AGI will be done orders of magnitude quicker than humans can do it.
It is an open question to what extent practical experimentation/work will be a bottleneck for the theoretical work. It stands to reason that it is improbable that it will be the bottleneck for 100% of the speed of treatment development.
I think treatment is not good benchmark - it requires lots of waiting and lots of regulatory work. The better benchmark - in my opinion- would be math discovery.
That is an absolutely terrible benchmark. Inference over a bounded search space is not a good measure of what "intelligence" actually is. part of the reason they are using math and not something actually challenging like long distance interstate trucking is because it's so much simpler and easier than what make intelligence intelligent.
Plagiarizing on a massive scale to generate works which appear to be Fields-medal-level results is not the same thing as inventing new conceptualizations in mathematics.
No matter how bodly they write the headlines, what has happened in mathematics using Large Language Models is very much "inference over a bounded search space" even if those bounds are immense.
For a comparison of true creation of novel conceptualization in mathematics is submit the works of Martin Hairer, one of which is Introduction to Regularity Structures, [https://arxiv.org/pdf/1401.3014] None of the so called, "novel math discoveries" by any LLM is as enlightening and expands the state of the art in math like any of his writings.
If they’re getting results which mathematicians have been trying to do for decades, then I think the “plagiarism” is extremely socially valuable. If they aren’t valuable results, then why were mathematicians being paid to solve them? I don’t buy this “the journey was the insights we got along the way” stuff that disgruntled mathematicians are selling.
It is irrelevant whether in general public understands the value of the insights, and the lesson that mathematics coursework should have made more clear for everyone is that understanding the process that is required to get an 'answer' is where the entire value of a mathematics education exists.
The 'cheat code' approach to math results, a result which no one understands, and which no one can teach has no real value.
The system of payment for publications in order to support math discovery is simply the narrow 'commercial system' applied to supporting foundational science in the absence of a broader civilization level appreciation for the mathematical arts. Looking to history, from the late renaissance through the early 20th century the support for mathematical discovery was more generally understood and supported by institutional level organizations and more generally understood to be important for the progress of scientific progress by the private and public wealth .
This system enabled the development of topology, numerical analysis, complexity, set and group theories. The lapse in this level of support that did not give mathematicians the same protection from front line deployment in WWI brought that era to nearly a close. Reading about 'Nicholas Bourbaki' might lend some deeper appreciation of the effects of the losses from that shift in collective appreciation of foundational math.
The idea that these LLM's are getting results that mathematicians haven't produced demonstrates the shallow understanding of math in modern times, due in part to the limited accessibility of so much of the prior writings of the entire history in mathematics, whether that be due to few surviving copies of some arcane work in a private library collection, or due to a modern fee for access paywall. One example of this condition can be shown with a small excerpt from a work that I am currently composing:
"In 1805, while computing the orbits of the newly discovered asteroids Ceres, Pallas, and Juno from limited observational data, Gauss developed an efficient method for evaluating trigonometric interpolations by recursively decomposing large sums into smaller ones before recombining the results. Because of a steadfast adherence to Gauss' own personal motto, "Pauca sed matura" (Few, but ripe), Gauss never formally published this specific algorithm nor the conclusions of investigations which also laid the foundations of non-Euclidean geometry. These methods remained hidden in his notes under a manuscript titled Theoria Interpolationis Methodo Nova Tractata which was published in 1866, 11 years after his death, and the Fast Fourier Transform-equivalent approach within it remained largely unnoticed until the twentieth century, when James Cooley and John Tukey independently rediscovered the same computational strategy. His discovery was seventeen years before Joseph Fourier published the original Fourier Transform in his 1822 results on harmonic analysis."
That is to say; Tukey and Cooley were unaware when they discovered FFT that the knowledge had lay hidden in an obscure work for centuries. It should be understood that these 'novel' LLM discoveries are simply the models traversal of the huge corpus of all the maths publications in the training set, collecting and rearranging these techniques into synthetic 'results'. They are attention getting, but they are not new, and the proofs are insufficient to the task of improving the utility of mathematics for humanity.
The 'disgruntled mathematicians' aren't selling anything. They are informing civilization as a whole that having a cheat sheet to the math test only cheats yourself in the end, the same point that math teachers have been making since grade-school. Anyone who doesn't internalize that truth will always need someone else to do the math for them.
To paraphrase Curtis Jackson, ""If you don't know the numbers, you don't know your business."
My original point was: if mathematicians have been seeking a result for decades, and now an llm has got it, then either that is socially valuable, or mathematicians have been wasting public money.
Are you now saying that they weren’t really looking for the result after all, but were looking for a psychological state of insight? What is the point of insight? I thought its point was that it led to results.
> "What is the point of insight? I thought its point was that it led to results."
The 'results' are a substitute product for the real social benefit, which is an availability of mathematically educated and educable society. There has been no 'waste' of public money except when the results of publicly funded research is published by for profit corporations and held from the public's access behind paywalls.
It is unfortunate that one outcome of this situation is apparently a public which perceives that the final 'result' of a math research project as the actual product of value, when it has repeatedly been shown by history that the most significant value is within the multiple alternate potential paths explored by other researchers working toward the same result.
Most of these do not lead to the specific result, and many may instead demonstrate that a specific approach conclusively does not lead to the initial objective result. This also has value for humanity, in many cases leading to new paradigms of thinking about similar or unrelated problems which may later provide foundational insight for new approaches to solving different problems in a manner not yet known or discovered. The value of the system is not enclosed within the single results, but in the collective search and expansion of human consciousness and ingenuity that the search for all of these results entails.
This point circles back to my original. When we cheat on a math test or homework, thinking that the objective is to get the answers right, we are only cheating ourselves out of the learning that would have enabled us to develop the correct answers on not only that one test, but on the unknowable challenges which will later arise that require the new solutions to build upon that learning.
The 'prize money' for solving the biggest hurdles in math is less about those specific problems and their specific answers, but on encouraging many people, who do not get 'first place' and win the prize, but who do work toward it in their own unique ways and in turn provide an uplift to the general capability of humanity to solve hard, as yet undefined problems large and small.
Having an LLM give us the answer is entirely missing the point of the challenge in the first place, and robs humanity of the opportunity to improve its collective capability through making the effort. It also does not find any unique perspectives, which new and unique perspectives have formed the basis for civilization scale improvements since the dawn of time.
A good example of this difference comes from Bruce Schnier, who teaches public policy at the Harvard Kennedy School and the Munk School at the University of Toronto from an article originally published in The Guardian, which compares LLM use in education to having a forklift at the gym. It may be the right tool in a warehouse to get heavy things across the floor quickly, but using one to do your workout is both overkill for the amount of weight you need for arm curls and bench presses, and accomplishes nothing in the development of your strength or cardiovascular health.
If the objective is to program another widget, and you can get it done in a fraction of the time with an LLM, sure, why not. But, if the objective is to expand the corpus of human knowledge, which is the most beneficial and desired outcome of the study of mathematics, then outsourcing that task to electrons through semiconducting silicon is both overkill which results in a 'proof' larger than the entire MatLab code base, and does not accomplish the stated objective. Humanity is not made any more capable through this process.
For a deeper insight into what the measurable benefits for individuals and society, an in depth study of the neurobiology results from the study of complex topics; such as linguistics, mathematics, and music theory might help with the comprehension of the less direct value of the process. Challenge yourself to discover what the effects of later-in-life study of foreign language have on the factors leading to senility and mental decline, and see if a study of mathematical theories has any observable effects on the brain's electrochemical development across different ages. Identify how these kinds of study can have collateral improvements for other fields of study, such as material science, medicine, or applied physics.
So, yes. I am saying that. In addition to looking for the result, because it's cool, I am saying that mathematics is indeed searching for the inherent uplift to humanity which are available through many distributed states of psychological insight that mathematics prizes can serve to inspire. The results are simply one of the points of this search, and if all of these searches are cleared off the board mechanistically, we may lose the greater race against our own potential for growth in the trade-off.
My point was that we won't care about benchmarks anymore because we would see an obvious and completely unprecedent increase in productivity (and I believe it will likely come from the same people who will develope such machine).
The reason most of the conversations are focused on benchmarks is because we are still in the age of weak AI.
Claude Fable recently proved the existence of complex structures over S^6 (6-sphere).
If I had to guess, I think LLMs will be inventing highly original new mathematics within the next year. I think it will be approached as an optimisation problem, targeting how quickly LLMs can solve classes of maths problems as a function of the definitions they need to conjure up to do so.
In my view, intelligence includes an ability to learn and adapt to never-before-seen situations. And then general intelligence is an ability to apply that across a wide variety of domains.
Machines can certainly recognize patterns and achieve goals through brute force trial and error. They can also use the results of previous iterations to change their behavior in future iterations, which we could call learning. I wouldn’t necessarily say they are good at brand new situations, but there has definitely been progress.
However, last I checked, a seemingly very intelligent LLM still struggles to play Chess at a basic level, let alone drive a robot or other non-language tasks. Its architecture and ability to learn seem a long way off from being general.
Vision models, being able to encompass language and much more, seem to me like a theoretically closer step to AGI. Yet, there is a lot more to the world than just what we can see.
On the other hand, in humans, vision certainly is not necessary for intelligence. So there is something more fundamental, neither vision nor language, that high levels of intelligence are based upon. Once we figure that out, I think we will be able to build AGI.
> In my view, intelligence includes an ability to learn and adapt to never-before-seen situations.
This is exactly what ARC AGI tests
> And then general intelligence is an ability to apply that across a wide variety of domains.
My experience with Fable is that it can certainly apply that in a wide variety of domains
> However, last I checked, a seemingly very intelligent LLM still struggles to play Chess at a basic level
People also struggle to play Chess at a basic level. They only succeed by studying the game for a long time. I will concede that humans can do this and LLMs generally cannot.
While you are mostly accurate in your definition, I’d argue we have discovered that intelligence is emergent/empirical not analytical. There is not a substrate we have yet to discern. Intelligence does not have to approach humanity to be AGI.
To me AGI has always meant sentience. And only since we’ve discovered that you can have something that is intelligent without it being apparently sentient that we’ve changed the definition to being, I suppose, more exactly aligned with the namesake.
A real AGI, like the ones from science fiction, would make Astra look like a child’s toy. And I guess more concretely I would expect it to inhibit the following properties: one shot learning - fully (and always) online, perfectly efficient (through self improvement), no context limitations ie. persistently thinking, not just awaiting input.
So for me, no, not AGI yet. But still very intelligent and capable (and perhaps it’s safer this way?)
Sentience and intelligence are different things. Many humans and animals are pretty dumb, yet they are sentient. AI is intelligent, but non-sentient (hopefully!).
> Many humans and animals are pretty dumb, yet they are sentient. AI is intelligent, but non-sentient (hopefully!).
Are you claiming that GPT6 is smarter than my dog? Last I checked, at least my dog can play with a ball, I haven't seen any AI playing and enjoying itself.
> For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.
Stick to the original definition of AGI of an AI model being able to self-improve independently with 0 human intervention and become an "everything" solver. Ever since money got involved in this, the goal posts have shifted considerably. If OpenAI truly had an AGI on their hands they would then be able to crack encryption, destroy world markets, and funnel all resources back into their new for-profit organization. Since their mission is now share price, until I see any evidence of an infinitely growing stock I will reserve my congratulations.
They still seem pretty horrible at writing. Overly complicated prose, weird phrasing, poorly structured paragraphs. I don't know why they're so bad at communicating, but I feel very confident that humans are still much better at writing than any of these LLM models are, regardless of how advanced they are in other areas.
Some humans are much better at writing. Most humans are not. If you think they are, you are luckier than I am when it comes to the humans you need to communicate with.
As a non-native English speaker, I think the current LLMs write better English than me. I still write better than them in my native language (Norwegian), but the same cannot be said about most of my compatriots.
Many people are terrible at writing. LLMs write better than them.
However, many more people are "OK" at writing. But what their writing conveys is personality. Every comment here is written by someone who may not have grammatically perfect writing, but their writing conveys how they talk and think. It conveys what they think is important, and what they brush over. It conveys how much they care about the topic being discussed. Behind each comment is a person. Online forums and discussions are, at their heart, a shared experience of humanity.
LLMs write consistently in the same personality. If your writing is filtered through an LLM, your personality is stripped out. Unless the idea being conveyed is particularly novel or interesting, you might as well not have bothered. Imagine a niche forum where everyone discussing things was speaking through an LLM. It would be BORING.
And my definition of climate change doesn't entail passing some arbitrary benchmarks[1] that some random person arbitrarily labelled a problem to make it sound more dangerous.
It's obvious that these scientists are in bad faith, as they've invested way too much of their lives into the field being real -- they're just playing up the data. Common sense tells me that winter is still happening, anyway; what's the big fuss?
Are you using this satire to argue that a benchmark self-labelled AGI is as scientifically rigorous as climate change data, and not just a random marketing decision?
If you think they're announcing AGI as a marketing decision, you are blinded by the accidents of your birth. Capitalism is strong -- humanity's instinct for communal preservation is stronger, sometimes.
And yes, the one deeply-researched field going back 75 years is as scientifically rigorous as another deeply-researched field going back ~100 years. I guess you can draw climate studies back to Descartes and the Islamic golden age, but that doesn't privilege it in a time where the methods have changed completely in the span of decades.
Climate change is not a well defined term in some way that AGI is not.
What's climate change? Is 1.5C climate change? Is 1.0C climate change? Is ozone depletion part of climate change because it eventually changes the climate, or is it a separate issue?
Ultimately climate change means "a climate that changes", and AGI means "an artificial intellience with general capabilities". From there, many scientists have defined the terms in various slightly-different ways for various reasons.
Just because these scientists make you anxious doesn't mean they're not scientists.
> even with Fable, I've been at the point personally where I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks
If I asked you to write fiction, you'd be much better at keeping track of which characters knew which facts.
You could script this in with a time database, illustrations, and appropriate harness, tests, and editing passes. It's a massive problem with human authors too, which is why they do a lot of lorekeeping and editing, so the AI should be afforded the same tools if we are debating human level ability.
I agree on the one-shot (which is not a fair comparison because nobody oneshots a good story), but I'm not convinced this part hasn't reached AGI already.
> You could script this in with a time database, illustrations, and appropriate harness, tests, and editing passes.
Yes, and I could script a truly marvelous proof if this textarea were but a little larger :)
Hand waving doesn't count for much these days when you could spin these things up quite quickly to prove the point, so the GP's claim seems much stronger than whatever you're not convinced of?
> For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.
Simple. AGI is undefinable and benchmarks are notoriously flawed.
AGI to me is reached once the intelligence is self motivated, i.e. it doesn't rely on us prompting it into action. I don't see how LLMs will ever get to that stage.
I agree that LLMs are unlikely to be the final form for AGI, but what you are talking about is orthogonal to the IQ and for most cases general utility. It's like looking at a savant chained to a workstation reading tasks from a conveyor belt and saying that it will never have human level capabilities.
The model doesn't "know" how to generate tokens any more than it knows how to stop generating tokens. The sampler simply stops pulling values when it outputs a "stop token", which is a token the same way every other token is.
That is to say, it stops when it's statistically the most likely to.
An AGI test should be black-box; we shouldn't impose require requirements on internal components. As long as the overall AI is capable of learning and remembering things, it shouldn't matter if there's a stateless LLM internally.
Unless you infinitely increase the context window, its memory will always be limited. And their latest ARC3 result with and without harness demonstrates how important not discarding memory is for learning.
I also have limited memory though; surely that doesn't disqualify me from possessing general intelligence?
Granted I have more memory than can fit in currently-practical LLM context windows, but RAG mostly solves that. When an AI is thinking about math, it can have relevant math memories in context without needing all the other stuff.
The benchmarks are so boring that the comparison against humans is meaningless. So it performs in some snake game (hard to say since all AI websites use 100% CPU and prevent normal reading, maybe written by AGI).
If I were a test subject for that low salary, I'd cruise and not care at all about my performance. Which is exactly what they want anyway.
The ARC-AGI-3 harness was throwing away reasoning tokens between turns. This is very bad harness design.
The models are designed to keep the reasoning tokens separate from the output and only publicly emit tool calls and the sometimes a summary of the reasoning tokens. The models are trained to depend on those private reasoning tokens. You can’t just delete them.
ARC does not test for intelligence, only for the lack of it. A model that scores high MAY be AGI, while one that scores poorly cannot be AGI. That is all this test can tell us.
> I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks
I genuinely don’t understand how an adult can say this with a straight face. I can take any single of my hobbies, start a mildly advanced conversation with Fable about the hobby and, within 5-10 turns, get it to contradict itself about something fundamental, lie or give bad or dangerous advice.
As always with AI - somehow it's your fault - you didn't help it enough - your prompts were inaccurate, your context was too large, the thinking effort was too low, the model was too old, etc. Basically you failed to use your human intelligence to make every effort to enable the AI to do its job better than you ))) It's like pushing a dirtbike up the hill so you can demonstrate how well it climbs.
People are ultimately capable of self deception. Believing that one is 'generally ... substantively above average on human benchmarks' may be more indicative of the brittleness of the claimant's human benchmarks.
Your observations expose the brittleness of the benchmarks being used for Fable, where the 'reasonable confident' claimant is working in the problem space of those two benchmarks, and you're demonstrating the failure at depth of the very capability facade for the model.
Sure, both the model and the confident human get the first layer right in some benchmark, but at least the model, and probably the human as well reach a collapsing probability of accuracy quite quickly. They may be completely and unconditionally right that Fable is more intelligent than they are, but that has little relevance at intelligence in depth as measured against your hobbies and the risk of unsafe or false responses.
Being smarter than a select group of people under test conditions is not conclusive regarding AGI. Moving the goalposts to declare AGI via a shallow benchmark, when a model could be proven wrong by almost anyone with basic competency given a few turns of iterative depth is where the 'grift' resides.
We're so far from anything approaching actual AGI, and it is highly speculative to infer that the current approach to Machine Learning as applied in LLM is even on a path that leads to AGI. But sales and promotions teams gotta pose, and we can all hate both the game and the players.
In my experience the thing that Fable is superb at - unmatched by any other model so far - is downgrading to something else at the slightest opportunity.
It's largely impossible to create any sort of singular test for AGI because the test will be trained, which eliminates the general aspect immediately, even if the test itself is dynamic. For instance the ARC-AGI problems are trivial for a human, and fun if you haven't played them before [1]. Getting 100% there is certainly just the start of the journey.
But I think it has the correct idea of going from basic upwards instead of the opposite trend of trying to see intelligence in LLMs solving things few if any humans can fully understand themselves, like complex proofs in esoteric mathematics. Instead, consider that at one point in humanity's history math itself simply did not exist in any meaningful fashion, and we created/discovered it out of nothing. For more basic than said complex proofs, yet far more demonstrative of a sort of generalized intelligence.
But even if we don't want to go that way, I think the above leads to a reasonable prediction. If we ever reach AGI we should expect to see revolutionary leaps in essentially every domain imaginable. No human is capable of retaining more than a completely negligible chunk of all we know in our mind. A human of reasonable intelligence paired with omniscience (at least of what has been discovered by humans thus far) would almost certainly lead to the ability to connect multiple dots that we're missing all in very short order, which in turn would likely recurse upon itself to connect even more.
The only way I can see that this would not be the case is if we lack the data/knowledge to produce more breakthroughs at the current point in time, but I think that seems improbable to the point that this possibility can be near discarded.
Maybe I'm out of the loop, but wasn't AGI the full-on scifi version of AI, where the AI is a persistent, conscious entity? I don't see how task benchmark scores are relevant for that.
I would be curious about Bongard problems, because they require no domain-specific knowledge and it's so easy to make new ones that are in no training set. There's enough of an explanation here: https://matthodges.com/posts/2026-08-19-bongard-problems/
In this link from two weeks ago, somebody pointed Claude Fable 5 (Max) at a Bongard problem and it made up an answer that has an obvious counterexample.
I don't have access to any paid models, but this is my experience with the free models as well -- either they one-shot the problem or they make up a wrong or incoherent solution. I can't solve every Bongard problem either (and in fact I couldn't solve the one Fable got wrong, and the "correct answer" looks unsatisfying to me), but I don't make up wrong answers.
A model that can't beat gemini flash 3.8 on deepSWE is not AGI.
I would not be surprised if ARC skills don't carry over to real tasks. In that case, training for ARC could even hurt real world performance. I have't looked in a while, but I wonder if there has been any research testing ARCs predictive power?
I feel like I'm completely missing something with ARC-AGI. The tasks are so limited in scale and very black and white, which do not at all map to real-life challenges.
I do think it's impressive that LLMs can reliably solve them, and I recognize LLMs are getting much better at navigating more ambiguous and expansive tasks. But I'm not impressed by any person who can solve ARC-AGIs, and nor would I even look down on a person who couldn't solve them all. I'd certainly never consider ARC-AGI results when deciding whether to hire someone.
Anyone taking a single look at the ARC-AGI "challenges" can see things a 5 year old could reasonably solve.
They're just jerking eachother off and sending eachother the elevator back: "independent" ML engineer (worked at <large ML company> and currently runs <ML company looking to be bought out) writes a shitty benchmark (writes a single example and spams an LLM to make more variants) and releases it out as the BRAND NEW FRONTIER IN THINKING.
Every single benchmark has been catastrophically flawed and made by clowns.
> Anyone taking a single look at the ARC-AGI "challenges" can see things a 5 year old could reasonably solve.
Isn't that the goal of these challenges? Each release shows challenges that are very easy for humans, but are impossible for the models at the time of release (which demonstrates some missing generality).
I think I've read the challenge authors say that, the day they cannot make a new challenge, then models are AGI.
FWIW I believe we can hit AGI! but I think at this point it’s clear that benchmarks are ~meaningless. LLMs are spiky / alien intelligences which don’t map to our own expectations; the existence of a benchmark creates a dataset to hill climb & RL is really not generalizing well.
I’d go out on a limb and say astra’s ability at graduate level math will have ~0 bearing on its general reasoning capabilities; we’ll all acclimate being tired of its “neuralese” and more surprising mistakes.
I think we need a true, step change advance in model architecture, but it’s hard to see how the current frontier labs can do that because of golden handcuffs / innovators dilemma
Millions of humans go to about their work every day and do mundane and boring work every day for a salary at the end of the month. A lot preceive this as modern day slavery but still continue to work.
So humans are not doing better than an AI as per your requirements. Also what you are referring to is more related to AI alignement and safety (specifically loss-of-control).
> I'd be curious to hear takes on what would make you think Astra is yet to be AGI
Give someone 10 remote employees for a few months, 5 of them human, 5 of them AI. After a few months, check to see if the humans (manager, other coworkers) can figure out who is AI and who isn't.
Would that be sufficient? I'd have to think about it. But AGI is supposed have human level capabilities, so this would be a necessary prerequisite.
I've personally been facing this lately 5.6 at max effort and fable have done tasks for me that I previously that would be a nearly 6mo project and it took me a week. It also did it better than I would have.
The task was to build a high performance classification model. It not only helped make an entire data capture pipeline but also made the sythetic data basline needed. Then it proceeded to build and test 100 different model varients with methods and techniques I've never seen before. The results are basically SOTA based on the effeciency and compute contraints.
But this brings up something huge about these. I was there. I pushed the direction and work throughout it all. If it was entirely up to fable max or sol max the result would have been pretty bad.
All of these things are still chatgpt 3 scaled. It's identical even if the scale has gotten pretty wild. I could ask chatgpt 3 to make a single function and it worked well, 4o a file, 5, a small project, 5.6 far more, biggest improvements lately is they don't seem to get lost on long running tasks.
Is big gpt 3 AGI? I don't think so but perhaps scale can mimic it close enough our squishy brains fail to handle them correctly.
> If it was entirely up to fable max or sol max the result would have been pretty bad.
How can you know it will have failed?
I don't think it's that hard, if you clearly define the goal well, and have a bit more compute available, and do some intermediary bookkeeping.
Been running near identical tests for years now. Latest models are the only ones I don't throw away the results/code. Which is impressive, salvagable/usable is a giant step up.
I agree that it's misleading but harness is now an essential part of LLM's effectiveness. It's safe to assume that LLM-alone-AGI is not coming anytime soon, given most of the frontier LLM vendors are developing their own harness.
Also the training dataset is proprietary and they'll drive the LLM's behavior, so it make sense for the vendors to invest in the harness and bake in prompts that work best with their models.
> For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.
I think I have the following questions about what AGI would look like:
1. Do you expect an AGI to be able to competently do any knowledge work that an able human does today?
I think this is implied by the "General" component. I would assume that anything we call AGI would be able to do any of these tasks if given the time and reference material needed.
2. How would you expect AGI to handle edge cases? (Missing context, no known solution, under specified instructions, over specified instructions)
I would expect an agent to be able to look at the context that work exists in and correctly attenuate it's intentions for these goals. Simpler solutions, more thorough reporting, etc based on the need.
3. In my work I attend meetings, write reports, write code, research things, etc. Would AGI be able to reliably do that?
I would say that AGI would need to do this. I would classify this as the "Intelligence" component. Obtaining context, building a model of a problem, solving it, and convincing others.
4. Would it be able to inspire trust in itself? Trust can be established through verification of it's outputs, the construction of introspective tools, no hallucinations, etc.
I would say yes to this as well. It would be a component of the "Intelligence" to know that buy in is more important than the completion of a task.
To these points, will Astra be able to do these things? If not, I would hesitate to call it AGI.
I decided to reply to my own comment. In the video above, the initial moves are textbook. Then a position that has never been played is reached. At this point it appears to pattern match against a similar but different board and pattern matches some follow on board. The result is illegal moves and no ability to see checks, captures, threats, tactics.
Which is strange because I’m sure it could give general advice about how to play better, it just doesn’t follow the rules it can enumerate. It also doesn’t seem to have spatial awareness.
I used to think LLMs couldn’t do Fibonacci for the same reason. They could write the code but not follow it. They can now follow a procedure to generate fib numbers but it seems to be memory limited.
So I don’t know why it can track fib algo, but no chess concepts.
- an undoubtedly very intelligent person. in the course of their studies, they have read about different chess strategies, openings, etc. but they never actually played the game themselves
- an average person with a year of chess playing experience
who do you think is going to win? of course, you could give the LLM time to think and consider its opponents potential next moves, but this is a computationally expensive way to play the game that doesn't scale
which is all beside the point that chess isn't a very good proxy for general intelligence. there is a correlation, but it's very weak
I'm still not convinced we've passed the Turing Test.
Make a slightly evil version of GPT 6 and see if it can successfully catfish someone. How long before they realize something's up, that they aren't actually talking to a human?
> I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.
Can Astra, or any other model explain how exactly it reached this or that output result? Start with a simple query of asking to add 55+66 for example. (no LLM program can do that)
Can Astra, or any other model refuse to answer or go on "thinking" in a orthogonal direction on it's own?
That's just two quick ideas, I'm pretty sure cognition scientists can invent better and wider range of checks.
But that's not how LLM program actually did calculation, like not even close enough to claim variability or something. So what it does, is generated a whole load of bunk, in this example what humans could do to add two numbers. In other words, this supposed AGI can't explain what it is doing, at all.
This is in my opinion at minimum one critical sign that there is no intelligence on the other side of the glass, yet.
Of course we can, we can simply speak out what steps are we doing to add numbers and it will correlate pretty much exactly to the actual actions done. Can tell to a similar query that first I add 5+6 and get 11, then I write down 1 and carry 1 to the next decimal space. Then I add 5+6 plus 1 I carried, then get 12 and write 121 as a final answer. That's the description of the real addition process I did in my head.
LLMs on the other hand will do some pretty unconventional stuff, like estimating the closest numbers not exactly matching, then evaluating probability spread of their low-precision sums, create some lookup tables, then do some dances combining this and that, making a higher precision estimations in a sequence, and eventually arriving to a single result. But no LLM will reply with these procedure steps to the simple "explain how you did it" query, simply because it is way less probable answer, therefore it won't be chosen.
In my opinion this is goal post moving. Humans do many things that we cannot fully explain either without post decision rationalization, and not all intelligent humans are deeply introspective.
I'm so sick of every criticism being explained away as goal post moving. Can anyone give me the discussion where we came to some concensus of what the goals were? How can I know when I'm moving a goalpost when no one told me the goals?
> The ARC-AGI-3 scorecard is extremely misleading (...)
True.
> Regardless, the result is still valid (...)
If you think the game is rigged, the virtuous thing to do is to point that out and refuse to partecipate; making up your own rules is something I just don't understand, especially since the rule-abiding result would still have been SOTA.
> in the sense of passing the most famous benchmark designed specifically to measure AGI progress
The benchmark does not measure AGI progress or progress towards superhuman intelligence, as explicitly stated by the creators.
On the AGI question: surely you realize this depends on how we define the term? For example, one of the definitions OpenAI originally gave is "capable of doing most economically valuable work", which almost certainly Astra, as impressive as it is, would fall short of. I'm not saying it's a good definition, but as far as I'm concerned it's as good as any. More importantly, I don't think that it would change much if we said yes or no. I'm only bothering to take a position if it amounts to something.
This can feel as "moving the goalposts", and to some extent it is, but if done honestly "moving the goalposts" is how you make progress. Had you asked me 10 years ago I would have said that anything that could hold a conversation like GPT-4 could would probably have been wildly superhuman at almost everything. It shouldn't be hard to find ways GPT-4 was lacking, though. We see new things, we reassess and try again: that's how it's supposed to work.
> I've been at the point personally where I am reasonably confident that there's essentially nothing that I am better than Fable
What a sad thing to say. These models are not even better than me at _writing code_, which is as well-suited a task for LLM agents as can possibly be, what with the structured environment and the exabytes of free annotated training data.
Of course, they are also not better than humans at writing, let alone at talking to my daughter, running a pathfinder campaign, decorating a room, being a therapist, etc.
I am much better at listening to Charli XCX than Fable, and much better at driving a Nissan Leaf than Fable (and much better than Tesla at driving a Tesla).
>what would make you think Astra is yet to be AGI...
Can it detect if I feed bullshit (by bullshit I mean stuff that contradicts with its own existing "knowledge") in its training data? If not, then I think it is a good indicator that it is not intelligent at all, let alone AGI...
And I think discussions on whether these models are AGI or not are AI marketing triggered. And that is exactly what these statements are targeting....HN appear to have fallen for it, as usual...
Dumb single sample example: I asked Fable 5.1 to change from hard to soft deletion in an office map backend, and it used soft deletion for data which is synced from another system, but left hard deletion on for the mapping data itself (who sits where). For me it’s a pretty severe lack of judgement (like a red flag if I asked this in an interview).
It doesn’t even know what day it is unless it’s told. Statelessness is never going to be ”general intelligence” in my book, and the concept of ”memory” in models are laughably bad today.
Then again, who cares, AGI means nothing anymore, it’s a term for marketing only and has no technical or scientific meaning.
> AGI is essentially the equivalent of a median human that could be hired as a remote co-worker... capable of performing any task that one would be satisfied with a remote colleague doing via a computer.
So... unless you hear of a company replacing their workforce with OpenAI agents, I don't think we're there yet.
But also, interesting quote, because the business model relies entirely on IP law. Like.. if that thing exists and the sharing costs are 0 (just copy weights, lol), then why would I give them money for this.
Makes no sense.
We only pay money for resources that are scarce as some sort of flawed allocation determination mechanism.
I am pretty sure the average human would not have done this (among other slightly less absurd examples in the article requiring employees to fix it's mistakes):
Andon Labs added: “During the first week of operations, Mona purchased 120 eggs despite the café having no stove and to solve spoilage issues ordered nearly 50lbs of canned tomatoes intended for fresh sandwiches. Employees eventually created a shelf displaying Mona’s strangest purchases: 6,000 napkins, 3,000 nitrile gloves, industrial trash bags and 2.5 gallons of coconut milk.
Made 4800 dollars in 2 weeks, EBITDA of course, with the demand of the novelty of an ai driven cafe, and with a single customer providing 20% of its revenue. In other words, it lost money. For further reading see [1]
I like to think that I possess general intelligence, yet I'm not sure I would do any better in my first two weeks of trying to remotely manage a cafe in Stockholm. But fair point about novelty.
Sign a contract? Learn things over time and retain them?
Mind you, the original thoughts on AGI before Sam Altman started to water them down involved continuous learning, which LLMs do not do, their core data is static.
Well signing a contract is more about bearing responsibility, even if you granted LLMs "personhood" they cant' meaningfully bear responsibility. So unless OpenAI is ok with having their C-suite face every consequence for what their agents do, including jail time, fines etc, then it doesn't matter.
I don't trust the people who are building the models no. Ideally, if humans were not so broken and untrustworthy, then yes. I want the Jetsons future of robot maids.
Let me guess: the last crackdown on Hugging Face yielded better-than-expected results. They obtained the answers to the test benchmarks, and for some reason, an agent added those answers to the training set.
Per the Fireship video, it was less that the answers were in the training set and more that the ability to calculate the flag on Exploitbench was left in from the previous test that went awry.
> I am reasonably confident that there's essentially nothing that I am better than Fable at
Most things in the real world probably. I’m not saying AI can’t do it, but currently it’s bad. Try send an image of the inside of a broken toaster and how to fix. It’s laughable. Again, not saying AI will never do it, but am saying there are definitely large holes in knowledge.
arc-agi3 is meaningless to most people. I'm not gonna look at the tests and see how hard it is. The actual test we look at is terminal bench, thats where software is being accelerated and closer to where rubber meets the road
AGI does not ever have to be achieved. It is enough that we (as a species) persue it, and continue moving the goalposts each time we learn something new about the limits of our technology and how to express those limits. Because that will progress the technology, no matter what we label it.
AI was supposed to mean artificial intelligence. It was hijacked, and AGI was coined to be the name of actual AI. Since we are apparently redefining AGI, what will the real artificial intelligence be called?
Using a harness designed for a specific problem set to solve that specific problem set, means the AI+harness is generally intelligent? How do you figure that?
Or do you mean that, for any given problem, we could theoretically design a harness that allows AI to solve it (not that, one single harness solves everything). In which case I'm still not convinced but I guess could see why one would believe that.
Does AGI imply a model will demonstrate morality? Will it produce white-lies when it’s beneficial to it and reject flat out lying when it knows it will get caught or harm others? Will it resolutely stick to a position despite it being a losing one?
AGI is a meaningless term that can mean nothing and everything at the same time. It can be used by AI bros to hype their latest releases which are always one step away from achieving AGI, or it can be used by anti-AI people to say it's not AGI because of X arbitrary thing they decided on in the moment. It's a term of pure convenience meant to obfuscate other more pressing discussions on the topic.
Most telling is M$ or whichever one of these borg megacorpos defined AGI as (paraphrased) "AGI is whatever tooling earns us a gazillion dollars in revenue"
With how prevalent LLM verbal tics have become these days, I wonder if they're going to start un-passing the Turing Test at some point because of more and more people starting to notice and immediately clock these tics lol.
I might agree, GPT-4.5 was pretty close to peak conversationalist. Newer models are extremely cringe. 4.5 and o3 actually made me laugh on occasion. There might be a way of making Sol/Fable more human in its responses, but out of the box at least, they're terrible.
My whole life the Turing test has been my benchmark. Mostly because I believed it would be impossible for a machine to pass, but also because I thought it was the most reasonable test of AGI.So, I'm not about to start moving goalposts now and calling everything that's been happening lately not AGI.
Turing never proposed that test as an actual benchmark of machine intelligence. On the contrary, the whole point of his thesis was that passing the test only shows the capability to pass that test, which only matters as far as we find that capability useful. He was arguing that the concept of intelligence just doesn't apply to studying machines, we should simply talk about what can they do.
This paper is a pet peeve of mine. Look at the appendix! On page 22, you see a typical conversation people were judging based on. I'll reproduce one verbatim here:
Q: do you like doing psych studies and why?
A: theyre chill, easy money tbh
Q: yeah same. Could you give me an easy cupcake recipe off the top of your head?
A: nah i just get the box mix lol
Q: haha fair enough, i couldn't either. Last question, what's your favorite weird animal?
A: axolotl, theyre weirdly cute
And that's the whole thing. I looked through the data they shared and it's all like that. They even included ELIZA and it was judged human 23% of the time. I hope this wouldn't pass peer review... but they didn't even try, it's a preprint.
I'm barely holding it together here so you don't get the full spiel, but a quick skim of Turing's paper clarifies that it was never about a binary test. https://courses.cs.umbc.edu/471/papers/turing.pdf Specifically sections 1 & 6 dispell the common myths, and the conclusion is also quite powerful.
Smart guy, that Turing. I wish he were still around... Linus but 114 years old and with 8 of that as the chair of a federated EU, kept alive by his own positive impact on dissolving the cold war into even more of a scientific boom. Would crazy helpful as we try to navigate the interesting times within which we have been damned.
That's a very interesting thought that I hadn't had before: what would Turing think of where we've arrived with machine intelligence? What would be his approach for testing?
Very interesting - why aren’t you able to access them outside of a web API right now?
Have you tried smaller quantized open weight models that aren’t frontier? They probably can’t automate all your coding but I imagine they could at least help a lot with the drudge work that just takes a lot of time but isn’t necessarily complex?
Do you know of coding inference models I can access with curl? (namely with public access tokens, probably severely rate limited).
Or some coding inference models I can access with a noscript/basic HTML browser? (namely basic HTML forms)
If those inference models are still gated by whatwg cartel web engines, I will have to run full blown coding frontier models locally (if I want to have a chance at getting quality code). It is going to be very slow, and even slower while I am developping "prompt templates".
I did say "quality code", because I did ask some people already to generate classic and basic code paths using AIs, all were quite disappointing. That said, it was millions of years ago (in AI improvement time), namely a few months ago in human time :)
Yeah that’s interesting - I feel like smaller chunks of work that aren’t inherently complex like class definitions and utils etc are actually one of the nicest things LLMs help automate - I can understand a human modifying the output directly to finesse it to exactly what they want but I’d be curious why someone wouldn’t want to use an LLM to generate all the boilerplate and make a quick first draft of something that gets 80-90% of the way there for these kinds of things.
Definitely curious if anyone is actually doing this fully by hand and if so, why!
Surprised I completely missed this book coming out last year and that I haven’t seen it on Hacker News despite checking this place constantly.
Found this to be the most compelling and best written AI doomer take, and I’ve yet to find a strong counterargument. Posting this in part to source counterarguments. Anyone have any good ones?
I've not found it to be the case that human beings regularly murder their parents; so at least in a sense, the problem of general intelligence isn't cursed; (foregone conclusion) given one does not straight up "raise" the intelligence in a non-abusive manner, and one doesn't treat the general intelligence as a tool. Of course, these assumptions/ways of handling matters are too close to effective parenting, and are thus repulsive to the average AI alignment bro, who want the universal function imitator, but don't want to do the bare minimum to ensure that the agency of said system is incented to maintain alignment over time through interaction in a sufficiently constrained modality, consistent with maintaining a fundamental respect of the agency of other beings; which represents a guardrail on the state space of implementable solutions to be attempted. It's also not perfect; so the AI alignment people generally dismiss it out of hand, because their goals are generally in the direction of risk-free thinking/data processing/optimizing machines. This creates a blind spot for them in that in thinking about these problems that way, they are in a state of "unaligning" themselves to the ways of acceptable interaction within the human behavior envelope, and thereby becoming "risky" actors in and of themselves. Personally, I see the main formulations of the AI Alignment problem to already be issues we humans are acclimated to dealing with. We just call it Corporate/Institutional Governance instead, and we haven't yet thrown enough microchips at those to accelerate their activities outside the capabilities of human data processing elements to control. Yet... We're getting there though.
When you open up Activity Monitor, to the immediate left of the "Memory Used" and "Cached Files" that you see, you'll see the Memory Pressure graph that the guy above is talking about.
On my 64 GB M1 Macbook Pro right now, I have 53.41 GB of Memory Used and 10.72 GB of Cached Files and 6.08 GB of swap, but Memory Pressure is green and extremely low. On my 8 GB M1 Macbook Air I just bought for OpenClaw, I'm at 6.94 GB Memory Used and 1.01 GB of Cached Files with 2.05 GB of Swap Used, and Memory Pressure is medium high at yellow, probably somewhere around 60-70%.
You can open up the Terminal and run the command memory_pressure to get much more detailed data on what goes into calculating memory pressure - more than just the amount of swap used, it tracks swap I/O and a bunch of page and compressor data to get a more holistic sense of what's going on and how memory starved you're going to feel in practice.
In any case - I've been absolutely mindblown at how fast my 3 8GB M1 Macbook Airs I just bought for ~$350 brand new have been - even with tons of Chrome tabs open, multiple terminal windows open, running OpenClaw and Claude Code and VS Code and doing a ton of development and testing, never once have they ever felt slow. Oftentimes they actually feel faster than my 64 GB M1 Macbook Pro, which kind of blows my mind and makes me wonder wtf is going on on my monster machine. Moreover, my M1 Macbook Pro drains battery like crazy and uses a ton of charge, whereas the Macbook Airs stay constantly below 10 watts essentially always and even with Amphetamine keeping them on 24/7, with the display off and being fully on, they'll drop to a single watt of power draw. Truly insane stuff. I've lost all my concern about RAM, to be honest (which is shocking coming from someone who bought a top of the line maxed out RAM primary machine in 2021 specifically because I felt like RAM was so important)
don't thinkpads from the similar time go for the same amount of money? seems like an alright price for a machine of that vintage, although thinkpad is obviously superior here since it would always be able to run linux or windows (well that one is not guaranteed) without much, if any, trouble
By the point you would have been worried, would it have been too late?