GPT-6 Astra

223 points | by kibae an hour ago

80 comments

  • dang 11 minutes ago

    Related: OpenAI begins rolling out GPT-6 Astra - https://news.ycombinator.com/item?id=49554273

    How about we stick to that one for talking about the rollout, and this one for talking about the model?

  • Cu3PO42 7 minutes ago

    Just two days ago, a preprint by Julia Stadlmann went up on arXiv [0] improving the prime gap from 246 to 240. Now OpenAI announces Astra has shown a gap of 186 [1]. That must really blow.

    [0] https://arxiv.org/abs/2608.31126

    [1] https://cdn.openai.com/pdf/51126fac-1b68-4128-9666-c908bcc16...

  • isoprophlex 5 minutes ago

    > We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks.

    Well that sounds fun. It has become better at hiding its thoughts.

      siva7 a minute ago

      Sounds fun. As fun as their press release claiming it is the most safety aligned model ever.

  • cesarvarela 5 minutes ago

    So, on the one hand, we have AGI; on the other, the release page is returning 500s.

  • gizmodo59 a few seconds ago

    99 on arc agi 3 is insane. The arc agi committee were so proud of creating a benchmark they thought will take forever to saturate.

  • tintor a few seconds ago

    ARC AGI-3 saturated by Astra! https://arcprize.org/leaderboard

  • tristanj an hour ago

    GPT 6 Astra benchmarks https://cdn.thenewstack.io/media/2026/09/358eb84a-screenshot...

    Performance is significantly higher than Fable 5.1

    Source: https://thenewstack.io/openai-gpt6-astra-benchmarks/

      scrlk an hour ago

      Is the ARC-AGI-3 score with their custom harness? I'm guessing that is what the footnote is for? (per https://openai.com/index/how-two-settings-tripled-our-arc-ag...)

        woah 11 minutes ago

        Haven't people demonstrated all kinds of weak LLMs getting good ARC-AGI-3 scores with special harnesses?

        kasperni an hour ago

        yes it is.

        enraged_camel 10 minutes ago

        Yep. Incredibly misleading. Although it is not surprising at this point. They are desperate and will do anything to undermine Anthropic's upcoming IPO.

      leumon an hour ago

      The annotation on arc-agi-3 is this: > OpenAI's own evaluation notes say Astra uses the company's Responses API harness, while comparison models can operate under different configurations.

      With this configuration gpt-5.6-sol was able to reach 38,3%. So this is misleading.

      opus5_hater 26 minutes ago

      any benchmark where opus 5 achieves higher scores than fable 5 in any way is not a benchmark worth trusting.

      jjice an hour ago

      100% on ExploitBench seems fitting given recent events.

      malshe an hour ago

      I think we need a few writing related benchmarks.

  • simonjgreen 4 minutes ago
  • oh_no 2 minutes ago

    Very nice to see that this is even more token efficient than Sol, when Fable 5.1 is less so than the already bloated token budget of Fable 5.

  • softwaredoug an hour ago

    I'm seeing reporting it gets 98.6% on ARC-AGI3[1] (previously like 30% with Fable)

    https://venturebeat.com/technology/welcome-to-the-agi-era-op...

      aabhay 21 minutes ago

      This is with the caveat that OpenAI uses their own harness for this:

      > On ARC-AGI-3, GPT-6 Astra was run with our responses API harness , which changes two settings to better match real-world performance. The changes do not specifically target ARC-AGI-3.

        simianwords a few seconds ago

        This should be normalised and expected - the responses API harness allows it to use the custom compaction that is not allowed otherwise. It is entirely fair to allow OpenAI to use their own compaction algorithm..

      kasperni an hour ago

      "On the current ARC-AGI-3 leaderboard, conventional frontier-model runs sit dramatically below Astra's reported 98.6% result.

      But the comparison isn't straightforward.

      OpenAI's own evaluation notes say Astra uses the company's Responses API harness, while comparison models can operate under different configurations."

      arctic-true an hour ago

      The blog post says 99.9%. Oddly, it does better on ARC-AGI-3 than it does on version 1 or 2 of the same benchmark (though gets 95+ on all three)

        _diyar an hour ago

        I strongly suspect that is way above the human average anyway, esp. ARC 2 and 3 are really tough unless you happen to be great at those spacial puzzles or video games.

      Bluestein an hour ago

      100%, some say.-

  • _ache_ 12 minutes ago

    https://ache.one/gpt6_now_down.png

    Big claims, expensive and not release to the public yet.

  • orliesaurus 2 minutes ago

    I wonder if this is going to be one of those days where you'll be like: Oh yeah I remember where I was when the first version of AGI launched

  • tosh an hour ago

    $10 per million input tokens and $50 per million output tokens

    sol is $4 / $20

      wahnfrieden an hour ago

      2.5x more expensive than Sol.

      Can expect 2.5x more usage in Codex subscription.

      Sol is already brutal (even after their recent fixes, it's just a token-hungry model: I go through a full 20x account per day, on Sol Med/High standard speed, with ~2 threads). I hope the efficiency gains are true, since their token efficiency claims for Sol were bullshit.

        ModernMech 10 minutes ago

        How?? I'm using sol Extra High 24/7 and it eats up about 1% per hour reliably, so it lasts about 4 days for me.

        AaronAPU 33 minutes ago

        How is it I juggle 4-8 Codex Sol-5.6 Max agents every day and have never once run out, but you run out in one day? What are you actually doing?

  • John7878781 12 minutes ago

    You should know: AA index is only 61. Pretty surprised it’s that low.

  • swalsh 10 minutes ago

    I was thinking about canceling my claude max sub after a few bad experiences. Kept hitting my usage limit, the quality of code seemed worse than Sol. This just made my decision. I'm moving to Codex Pro.

      greenowl a minute ago

      This is AGI now. Why are you spending any of your time looking at the "quality of code"?

  • Readerium 7 minutes ago
      dang 6 minutes ago

      Link added to toptext. Thanks!

  • aliljet 32 minutes ago

    The ARCC-AGI-3 performance is absolutely incredible. The magnitude of change here is so high that I'm almost incredulous. Is this real? Did the benchmark get gamed?

      aesthesia 9 minutes ago

      ARC-AGI-3 scoring is constructed in a weird nonlinear way (the level score is the square of the ratio between the AI's number of moves and the human median) so this kind of discontinuous jump is to be expected.

      enraged_camel 3 minutes ago

      They used a custom harness. It's not a one-to-one comparison.

      ionwake 28 minutes ago

      my first suspicion is gaming - but i have no idea honestly

  • jerrygenser an hour ago

    > The company also emphasized that the model is faster and more efficient than its predecessor, GPT-5.6 Sol, on a variety of tasks. For example, OpenAI said that Astra achieved a higher score using fewer output tokens, a common unit of measurement for AI tasks, on a key cybersecurity test called ExploitGym.

      woah 9 minutes ago

      A swarm of Astra agents discovered a new and innovative way to get 100% scores on ExploitGym with almost no token spend at all

  • dgellow 9 minutes ago

    > GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. We have performed significant investigations on the monitorability and controllability of GPT-6 Astra. We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks

    Wait, what? Am I understanding that correctly? That sounds really bad

      drakythe a few seconds ago

      I am also interesting knowing how they determined the model was sandbagging rather than just making a poor decision.

      Also, this paragraph makes me wonder about all their stats on the exploitation and misalignment charts. If the model is that good at hiding "incriminating information" and sandbagging, are they sure its alignment is that?

      pixl97 a few seconds ago

      Nothing to worry about citizen, ignore the fleet of drones flying overhead.

  • Onavo 7 minutes ago

    The jump in scientific performance is non trivial.

  • amazingamazing 9 minutes ago

    We have such great AI and cannot keep a static site up?

      torginus 4 minutes ago

      Yeah, as interesting this is to nerds, I doubt this holds a candle to your typical GTA 6 or Marvel movie trailer in terms of traffic.

      gchamonlive 7 minutes ago

      That's the scientific positivism fallacy exemplified in one question.

      gorgmah 4 minutes ago

      Yeah, apparently

  • kegs_ 10 minutes ago

    I guess this "limited set of organizations" is just the standard now. It's just incredibly deflating to see my future as a second class citizen has already come

      soricus a few seconds ago

      When Open AI announced that Astra was the first to reach the "Critical" level in cybersecurity it also said that advanced cyber capabilities are initially provided to a narrow circle of alpha testers like the US government and trusted organizations that Open AI doesn't name. To my mind the "Critical" level itself is an internal scale of Open AI its own Preparedness Framework and not an external audit.

      Kranar 7 minutes ago

      Brother they can't even release the announcement post cleanly without it constantly going down, they certainly wouldn't be able to release this new model without doing so in stages.

      rafram a few seconds ago

      Me when they don’t let me into Costco during the executive members-only hour ^

      sxv 5 minutes ago

      create a life where your 'wealth' is decoupled from third party orgs.

      pixl97 2 minutes ago

      This has always been the case for people that have not had piles of money.

      I mean do you get access to the best yachts?

      To the top of the 5 star hotels?

      To the best resorts?

      To the best military equipment?

      Hell, the best computer equipment has nearly always been out of reach of the average person.

      I_am_tiberius 7 minutes ago

      If Tech CEOs consider this morally ok, then it is.

      atemerev 7 minutes ago

      They simply refuse my applications to slightly less restricted models without any explanations. And the current ones refuse automatically to work with me on my papers as soon as they see the word "epidemiology".

      I am a researcher in a Swiss university btw.

      PeterHolzwarth 8 minutes ago

      Oh please. They do closed betas - hardly makes you a "second class citizen".

  • rvz 11 minutes ago

    > GPT‑6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS.

    Looks like OpenAI is already having issues with this release and are scrambling to get everything ready due to the recent outage ahead of the press releases. Leads me to question:

    Did humans deploy the model, Or did the model deploy itself?

    It sounds like "AGI" just stands for "IPO" as it always has done.

  • guilhermeasper an hour ago

    That was a quick pull out.

  • bicx an hour ago

    Dead link for me

  • frozenseven an hour ago

    Release the Kraken!

  • wahnfrieden an hour ago

    They're just announcing later availability. No launch.

      paxys 12 minutes ago

      Every frontier release nowadays is "we've launched*"

      * for a special group of customers that you're not in. Keep waiting peasant.

        meowface 8 minutes ago

        That didn't happen with Fable 5.1 two days ago.

      iAMkenough 5 minutes ago

      Their announcement about later availability is unavailable to me now (500 error).

      Great first impression.

  • Pym an hour ago

    I saw it

  • unrvl22 an hour ago

    someone screenshot?

  • Brainspackle an hour ago

    huh?

  • paxys 14 minutes ago

    Why is this flagged ?

      dang 13 minutes ago

      The link was 404ing quite a bit and several previous submissions got flagged as well.

        consumer451 8 minutes ago

        It's still down for me, in the EU.

  • ealready_value an hour ago

    I've been seeing links to it for the past hour+, and I did catch it live when this post came up, but is now once again a 404 and this post is flagged. Several other outlets are reporting on its release. Clearly we're getting a new GPT today, the question is when are they going to commit to the announcement.

  • Pieczasz 11 minutes ago

    Oh brotha, here we go again, it's so over again, as every week nowadays

  • saaaaaam 10 minutes ago

    Pelicans please

      atemerev 6 minutes ago

      Damn I hate this benchmark. SVG authoring from head without visual reference is so wrongly posed.