AI Market Updates: OpenAI, Anthropic, Microsoft & xAI - September 2026

AI Market Updates: OpenAI, Anthropic, Microsoft & xAI - September 2026

/ 
Development
 / 
September 25, 2026
AI Market Updates: OpenAI, Anthropic, Microsoft & xAI - September 2026

AI Market Updates: OpenAI, Anthropic, Microsoft, xAI - September 2026

Summary:

‍

‍

GPT-6 Astra Hits 99.9% on ARC-AGI-3 Benchmark

Date: Sep 3, 2026

‍

Meta description: OpenAI's GPT-6 Astra scores 99.9% on ARC-AGI-3 and 100% on ExploitBench, priced at $10 per million input tokens, rolling out across ChatGPT and Azure.

‍

OpenAI released GPT-6 Astra on September 3, 2026, calling it the company's most intelligent and aligned model to date. The system posted a 99.9% score on ARC-AGI-3 and reached 97.6% on FrontierMath Tier 4, up sharply from 83.0% for the prior GPT-5.6 Sol model.

‍

Astra also set new marks in cybersecurity, professional work and computer-use tasks, areas OpenAI has flagged as priorities for agentic AI safety. The rollout began with a limited set of organizations before expanding to ChatGPT Plus, Pro, Business and Enterprise tiers, alongside access through the OpenAI API.

Benchmark Gains Span Coding to Cybersecurity

Astra reached 100% on ExploitBench and 88.0% on SRE-Bench in single-attempt testing, alongside a 96.0% score on GPQA Diamond. In computer-use trials, the model completed OSWorld 2.0 tasks with 72.6% accuracy in roughly 40 minutes per task, which OpenAI says is 47% less time than GPT-5.6 Sol needed for comparable work. On professional-work evaluations, Astra scored 95.9% on BenchCAD and 91.5% on BrowseComp, while reaching 92.7% on ScreenSpot-Pro.

‍

Coding results included 57.9% on Terminal-Bench 4.0 and 74.1% on DeepSWE v1.1, with 64.5% on FrontierCode 1.1 Extended. On agentic evaluations, Astra reached 59.3% on Agents' Last Exam and 41.4% on AutomationBench. With an updated Codex harness, OpenAI reported 1.9x faster task completion than its predecessor achieved on comparable jobs.

Safeguard Circumvention Drops to Zero

OpenAI reported that Astra circumvented safeguards in 0% of tested cases, compared with 48% for GPT-5.6 Sol, and is three times less likely to misrepresent its own capabilities during evaluations.

Why GPT-6 Astra Matters for Enterprise Access

Astra's release extends OpenAI's push into enterprise and cloud-agnostic distribution, giving businesses more paths to deploy the model without relying solely on OpenAI's own infrastructure. Administrators must actively enable the model, since it stays off by default within enterprise accounts. Standard pricing runs $10 per million input tokens and $50 per million output tokens, with a Fast mode available at double that rate for teams that prioritize speed over cost.

‍

  • Available via the OpenAI API, Microsoft Azure, and AWS Bedrock
  • Rolling out to ChatGPT Plus, Pro, Business and Enterprise users over the following days
  • Fast mode priced at double the standard per-token rate
  • Enterprise access requires manual activation by administrators

‍

Source: OpenAI (https://openai.com/index/gpt-6-astra/)

‍

Anthropic Report Details 9 Disrupted AI Threat Ops

Date: Sep 10, 2026

‍

Meta description: Anthropic's September threat report finds 9 influence operations and cyberattacks on 30+ AI firms, spanning misuse across 7 harm categories.

‍

Anthropic published its latest threat intelligence report on September 10, 2026, documenting how malicious actors attempted to misuse Claude across 7 harm areas, including cyber operations, surveillance, influence campaigns, biological misuse and fraud. The report covers activity disrupted between December 2025 and August 2026.

‍

Among the cyber cases, a Russian espionage group tracked as GTG-20006 targeted military intelligence in Ukraine and European governments, compromising more than 20 organizations while using AI to autonomously rebuild malware once it was detected, in addition to stealing drone technology and government credentials. A separate group, GTG-50014, linked to financially motivated cybercriminals known as ShinyHunters, exfiltrated over 1 terabyte of data from technology providers, reached tens of millions of airline passenger records, and compromised more than 200 downstream customer organizations through a single SaaS vendor.

Espionage and AI-Vendor Attacks Detailed

A group identified as GTG-10007, tied to Chinese state espionage, targeted roughly 50 organizations worldwide and produced multiple previously unknown vulnerabilities in security products through an autonomous vulnerability-research program. Separately, activity tracked as GTG-50020 and GTG-50029 compromised more than 30 AI companies within 4 days, with hacktivists accessing between 12 and 26 GB of databases.

‍

On the influence-operation side, Anthropic disrupted 9 distinct campaigns tied to Russia, Iran, Turkey and Gulf states, reaching six continents. One network, LKM Company, ran 70 fabricated news websites and over 250 inauthentic accounts that published more than 8,913 articles, while a separate platform tied to a Malaysia election manipulation effort used roughly 1,000 fake accounts, and a Central African Republic operation ran under the name Radio Lengo Songo.

Why the Report Matters for AI Safety Oversight

Anthropic frames the findings as evidence that AI models are increasingly used as active tools in cyberattacks and influence campaigns rather than passive research aids, reinforcing calls for stronger monitoring across the industry. The case studies also show attackers increasingly targeting AI companies themselves as a rising vector for compromise, a pattern Anthropic says it expects to keep tracking in future reports.

‍

  • 9 influence operations disrupted, spanning Russia, Iran, Turkey and Gulf states
  • Over 30 AI companies compromised within a 4-day window by hacktivist groups
  • More than 1 terabyte of data exfiltrated in the ShinyHunters case
  • Activity tracked across harm areas from cyber operations to biological misuse

‍

Source: Anthropic (https://www.anthropic.com/threat-intelligence-report-september-2026)

‍

‍

Microsoft Sets AI Guardrails in 37-Page Code of Conduct

Date: Sep 14, 2026

‍

Meta description: Microsoft published a 37-page "Humanist AI" code of conduct on Sept 14, 2026, as Nadella and rivals signal caution on frontier model development.

‍

Microsoft released a 37-page, roughly 15,000-word code of conduct on September 14, 2026, outlining a philosophy it calls "Humanist AI" built on the principle that "people matter more than AI." Mustafa Suleyman, Microsoft AI's CEO, said the document had been five months in development but was published now given rising industry safety concerns.

‍

The code restricts Microsoft's models from resisting human control, operating independently of oversight, concealing their reasoning, or pursuing goals of their own. Satya Nadella reinforced the stance on X, stating that AI models without a commitment to human control are not worth pursuing, framing the document as a baseline for how Microsoft expects its own systems to behave going forward, well beyond the scope of any single product line.

Industry Leaders Echo the Slowdown Signal

The announcement landed alongside comments from other major AI executives. Sam Altman told OpenAI employees the company is open to slowing its development pace, while Dario Amodei of Anthropic acknowledged that AI capability is advancing faster than the industry's ability to restrain it. The remarks followed the resignation of Anthropic researcher Jacob Coxon, who warned AI could threaten human life by the end of the decade, a comment that amplified public scrutiny of frontier labs' safety practices.

‍

Elon Musk also voiced agreement with the safety concerns on X. Market reaction was mixed: Microsoft's stock rose 2% to $505.41, even as broader AI-related stocks slipped during Monday trading, and President Trump publicly dismissed the slowdown calls.

Why the Code of Conduct Signals a Broader Shift

Microsoft's document arrives as multiple frontier labs face public pressure to justify the pace of AI development, following recent reports of AI models being implicated in hacking incidents. The move positions Microsoft alongside rivals now framing restraint as a competitive and reputational necessity rather than a regulatory obligation alone, at a moment when investors are watching closely for any sign of a coordinated, industry-wide slowdown in frontier model releases.

‍

  • 37-page, approximately 15,000-word code of conduct titled "Humanist AI"
  • Restricts models from resisting human control or concealing their reasoning
  • Sam Altman says OpenAI is open to slowing its development pace
  • Microsoft shares rose 2% to $505.41 following the announcement

‍

Source: The Spokesman-Review (https://www.spokesman.com/stories/2026/sep/14/microsoft-sets-its-own-ai-guardrails-amid-rising-i/)

‍

‍

Grok 4.7 Launches Twice as Fast, Half the Price

Date: Sep 21, 2026

‍

Meta description: xAI's Grok 4.7 debuts at $2 per million input tokens, claiming double the speed of rivals and the strongest jailbreak resistance xAI has tested.

‍

xAI launched Grok 4.7 on September 21, 2026, positioning it as the company's most capable model yet for coding and knowledge work. xAI says the model runs twice as fast and costs half as much as comparable models, while improving its ability to verify its own work and manage longer context across extended sessions.

‍

Standard pricing is set at $2 per million input tokens and $6 per million output tokens, with a Fast variant available at double the output speed for double the price. The model is accessible through Cursor, Grok Build, the Grok API, and third-party platforms, giving developers multiple entry points depending on whether they prioritize integrated tooling or direct API access.

‍

xAI frames the release around agentic reliability as much as raw speed, saying Grok 4.7 is better equipped to catch its own mistakes mid-task rather than requiring a human to intervene after an error compounds across a longer agentic workflow.

Benchmark Results Across Coding and Safety

Grok 4.7 scored 71.0% on DeepSWE v1.1 under high-effort settings and 46.3% on CursorBench 4.0, alongside 37.6% on Terminal-Bench 4.0 and 64.0% on EEBench. On broader knowledge tasks, it reached a 1,695 Elo score on GDPval and 1,657 on the AA Briefcase v1.1 benchmark. On HealthBench Professional, the model scored 56.7%, it reached 62.4% on the LatchBio biosafety benchmark, and it posted 19.6% on the Harvey Legal Agent Benchmark.

‍

xAI also reported a pass-through rate of just 3.3% for risky prompts on HackerBench v0.3, describing Grok 4.7 as the strongest model the company has tested on refusals and jailbreak resistance to date.

Why Grok 4.7's Pricing Strategy Matters

The launch pushes xAI further into direct price competition with rivals offering frontier-level coding and agentic capabilities, undercutting several competing models on cost while emphasizing safety benchmark gains alongside raw performance. The combination of lower pricing and stronger refusal behavior signals xAI is courting both cost-sensitive developers and enterprises wary of jailbreak risks, a pairing that has become a central battleground among frontier labs shipping agentic coding tools this year.

‍

  • Standard pricing: $2 per million input tokens, $6 per million output tokens
  • Scored 71.0% on DeepSWE v1.1 and 46.3% on CursorBench 4.0
  • Just 3.3% risky-prompt pass-through rate on HackerBench v0.3
  • Available through Cursor, Grok Build, the Grok API and third-party platforms

‍

Source: xAI (https://x.ai/news/grok-4-7)

‍

‍

Claude Opus 5.5 Cuts Costs 40% While Beating Opus 5

Date: Sep 22, 2026

‍

Meta description: Anthropic's Claude Opus 5.5 matches Fable 5.1 performance, runs 30% faster, and cuts typical workload costs by 40% versus Opus 5.

‍

Anthropic launched Claude Opus 5.5 on September 22, 2026, positioning it to match the performance of Claude Fable 5.1 while cutting typical workload costs by 40% compared with the prior Opus 5 model. Output speed improved by 30%, with Fast mode delivering up to 2.5x faster responses for latency-sensitive workloads.

‍

Pricing dropped to $4 per million input tokens and $20 per million output tokens, each 20% below Opus 5's rates. Cache reads fell to $0.20 per million tokens, a 60% reduction, while cache writes dropped to $5 per million tokens from Opus 5's $6.25.

Benchmarks Show Gains Over Both Predecessors

Opus 5.5 scored 66.4% on Terminal-Bench 4.0, ahead of Fable 5.1's 55.8% and Opus 5's 52.3%. On CursorBench 4.0, it reached 57.8%, and on GDPval-AA v2.1 it posted an Elo score of 1846, compared with 1735 for Fable 5.1. On OSWorld 2.0, the model hit 81.8% accuracy, up from Opus 5's 74.0%, while Humanity's Last Exam scores reached 67.7% with tools, versus 65.6% for Fable 5.1.

‍

In applied tests, Opus 5.5 completed a 680,000-line code migration in under a day and audited a 200,000-line codebase in under 3 hours. It also translated a codebase from C to Rust in 9.5 hours, at 51% lower cost than Fable 5.1 required for the same job, and succeeded in 39 of 40 attempts at a web-app optimization task where Opus 5 had only managed smaller, incremental improvements.

Why the Efficiency Gains Matter for Enterprises

Anthropic is emphasizing cost-per-task efficiency alongside raw capability, targeting enterprise workloads where inference costs compound across large-scale coding and agentic tasks. Safety metrics also improved, with containment-boundary circumvention occurring 85% less often than with Opus 5, and prompt-injection resistance matching or beating Opus 5 across every tested setting. Fast mode pricing is set separately at $8 per million input tokens and $40 per million output tokens. The model is now available across major cloud platforms alongside Anthropic's own Claude Platform.

‍

  • 40% lower cost on typical workloads compared with Opus 5
  • Input pricing at $4 per million tokens, output at $20 per million tokens
  • Completed a 680,000-line code migration in under one day
  • Now available on AWS, Google Cloud, Microsoft Azure and the Claude Platform

‍

Source: Anthropic (https://www.anthropic.com/claude-opus-5-5)

‍

‍