The Problems Behind The Astra Vs Fable Benchmark’s Point Reduction
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: The Problems Behind The Astra Vs Fable Benchmark’s Point Reduction on ThorstenMeyerAI.com

TL;DR

Recent analysis shows the Astra vs Fable benchmark’s reported point reduction stems from index revisions, architectural changes, and misinterpretations. The true story is more nuanced, affecting how AI performance is assessed and compared.

Recent scrutiny of the Astra vs Fable benchmark reveals that the reported five-point drop in scores is primarily due to index revisions and architectural shifts, not a decline in AI performance. This development impacts how AI capabilities are compared and understood, especially in terms of cost-efficiency versus raw intelligence.

Thorsten Meyer, who gained early access to GPT-6 Astra, analyzed the benchmark data and found discrepancies in the reported scores. The circulating comparison, which showed Astra outperforming Fable on the Artificial Analysis Intelligence Index with a five-point margin, was based on an outdated version of the index. After the index’s revision—moving from version 4.1.1 to 4.2—scores for both models shifted, reducing the apparent gap from five points to two.

Further, the initial narrative suggested Astra was more economically efficient than Fable, based on token usage and cost per task. However, Meyer’s analysis clarifies that the index’s scoring metrics are heavily influenced by architectural features. Astra’s recent design involves latent reasoning loops that do not produce tokens in the traditional sense, making token counts a poor proxy for actual compute or intelligence. Therefore, the cost and efficiency claims based on token metrics are misleading.

Artificial Analysis’s own evaluation indicates Astra is less cost-effective for general intelligence tasks compared to its predecessor, despite being more efficient for coding tasks due to reduced token usage. The confusion stems from conflating architecture-specific efficiencies with overall intelligence performance, which are measured differently across various indices.

At a glance
analysisWhen: developing; issues surfaced following A…
The developmentThe Astra vs Fable benchmark’s point reduction is caused by index revisions and architectural changes, leading to confusion over AI performance metrics.
Five Points That Became Two — Reality Check
AI Dispatch · Reality Check · 5 September 2026

Five points that became two: what’s wrong with the Astra vs Fable benchmark

The comparison everyone is quoting — Fable 66, Astra 61, “not a rounding error” — is built on numbers that were stale when written, measuring a quantity that no longer means what it used to, aggregated in a way that hides the reversals that matter. The benchmark isn’t broken. The way it’s being read is.

Problem 1 — the numbers moved: Index v4.1.1 → v4.2 during launch week
as quotedAA today (v4.2)
Claude Fable 5.16657 max effort
GPT-6 Astra6155 max · a third source says 60
gap: 5 2 “Five points is not a rounding error.” Two points, on an aggregate of ten evals that just swapped three of them (GPQA Diamond out; AA-Briefcase + GDP.pdf in), is exactly a rounding error. Not AA’s fault — revising an index is how you keep it honest. The error is downstream: quote the version, or don’t quote the number.
◆ Problem 3 — the tell: Astra’s effort dial isn’t connected to the engine
non-reasoning554.4M tok · $990
medium52
xhigh54$2,778
max55$3,020
Non-reasoning = max. Same score, 3× the cost. Because Astra is reported to be a looped / recurrent-depth transformer — it reasons in latent space, without emitting tokens. The Index prices cost in tokens, measures verbosity in tokens, computes time in tokens. For this architecture it’s counting the receipt, not the work. “140M vs 42M tokens” compares Fable’s verbalized reasoning to Astra’s post-loop output — an artefact, not an efficiency finding. Nobody outside OpenAI knows what the loops cost in GPU-seconds.
The other three problems
02
AA’s own conclusion is the opposite of the story
AA’s benchmarking note: Astra is 75% more expensive than GPT-5.6 Sol at max effort and “largely sits behind its predecessor on the Intelligence Index vs cost frontier.” Price went 2.5× ($4/$20 → $10/$50); token savings only partly offset it. The genuine efficiency win lives in one place: the Coding Agent Index, where Astra equals Fable 5 at under half the cost. “Astra attacks the economics” stretched a true coding result over an intelligence index where AA says the reverse.
04
“Max effort” isn’t the same experiment twice
Fable at max = more tokens. Astra at max = ~nothing (see ladder). And OpenAI’s docs say Astra does not support `none` reasoning effort — yet AA lists a “non-reasoning” score. The most efficient-looking config on the leaderboard may not be one you can buy.
05
The aggregate hides the reversals
Index: Fable +2. OpenAI’s own evals (self-reported): Astra ahead 6 of 7 — AutomationBench, BenchCAD, Terminal-Bench 4.0, DeepSWE, TB-Science, FrontierMath T4; Fable takes HLE+tools. A 6–1 task split became a two-point average, and the average became the story. Ten choices deep, two points is noise wearing a number.
✓ What actually changed — and it’s not on the leaderboard
Hallucination rate 92% → 51% on AA-Omniscience — a 41-point drop; matters more than any 2 Index points “Same headline price” hides cache read $1.00 vs $0.25 (4×) + a 25% cache-write premium — the line that dominates agentic bills Coding Agent Index: Astra = Fable 5 at < half the cost — real, and narrow
The take

Three things happened at once: the Index was revised (five became two), the architecture changed (tokens stopped being compute), and the aggregate did what aggregates do (6–1 became +2). A leaderboard position now tells you less than it ever has — and the more advanced the architecture, the less it tells you. Latent reasoning is only the first architecture to break the token proxy. So with your Astra access: ignore the Index number. Take your ten real tasks. Run both models at the effort setting you’ll actually pay for. Measure the bill including the cache line. Measure the failure rate — the 41-point hallucination drop is the one number here I’d bet money on. The benchmark can’t decide for you anymore.

Sources: Artificial Analysis Intelligence Index v4.2 / v4.1.1 model pages (Fable 5.1; Astra non-reasoning/low/medium/high/xhigh/max — scores, tokens, total cost, cost-per-task method incl. cache-write & reasoning tokens) and “Benchmarking GPT-6 Astra” (75% vs Sol, cost frontier, Coding Agent Index, 92%→51% hallucination, 2.5× price, cache terms); OpenAI GPT-6 Astra developer docs (`none` unsupported, cache-write billing, logprobs removed); Alan D. Thompson, The Memo 4 Sep 2026 (looped-transformer read, unconfirmed); OpenAI’s self-reported Astra-vs-Fable table; the circulating 66/61 comparison (pre-v4.2). Scores are version-dependent and were changing at time of writing. Not investment advice.
thorstenmeyerai.com

Implications for AI Performance Comparisons

This analysis underscores the risks of relying on static benchmark scores when models and indices are constantly evolving. Misinterpretation of data can lead to overestimating or underestimating an AI’s capabilities, affecting investment, development priorities, and public perception. For stakeholders, understanding the architectural nuances—such as Astra’s latent reasoning loops—is crucial for accurate evaluation.

Moreover, the misalignment between token-based metrics and actual compute or intelligence performance highlights a need for more transparent and architecture-aware benchmarking methods. As models become more complex, traditional proxies like token counts may no longer suffice, necessitating new standards for measuring AI efficiency and effectiveness.

Amazon

AI benchmarking analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Astra and Benchmark Revisions

GPT-6 Astra was launched with a new architecture that employs latent reasoning loops, allowing it to process more tasks without producing additional tokens. This design was not fully accounted for in existing benchmarks, which primarily measure output tokens and associated costs. The Artificial Analysis Intelligence Index, widely used to compare models, has undergone several revisions—moving from version 4.1.1 to 4.2—adding new evaluation metrics and dropping previous ones like GPQA Diamond.

These updates have caused scores to shift across models, creating confusion when comparing results over time. The initial reports, based on earlier index versions, suggested Astra was outperforming Fable in both intelligence and economics. However, subsequent revisions and deeper analysis reveal that these conclusions are not as clear-cut, especially given architectural differences and the limitations of token-based metrics.

Historically, benchmarks have struggled to keep pace with innovations in model design, leading to mismatched comparisons and overreliance on outdated data. The Astra case exemplifies this challenge, emphasizing the importance of context-aware evaluation methods.

“The five-point gap that ‘is not a rounding error’ is actually within the margin of error once index revisions are accounted for.”

— Thorsten Meyer

Building Robust AI Evals: Proven Strategies for Testing, Monitoring, and Improving LLM Performance (Engineered: Data, AI, and DevOps)

Building Robust AI Evals: Proven Strategies for Testing, Monitoring, and Improving LLM Performance (Engineered: Data, AI, and DevOps)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About Astra’s True Performance

It remains unclear how Astra’s architectural innovations will influence future benchmarking standards and whether new metrics will emerge to better capture its capabilities. OpenAI has not publicly detailed the full computational costs of latent reasoning loops, leaving questions about the true efficiency and performance benefits of Astra’s design.

Additionally, the impact of index revisions on historical comparisons raises concerns about the stability and reliability of benchmark scores over time. It is not yet confirmed whether future updates will stabilize these metrics or lead to further fluctuations.

Amazon

AI model comparison tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Benchmarking and Model Evaluation

Expect ongoing revisions to benchmarking indices as models evolve and new architectures like Astra become more prevalent. Researchers and evaluators will need to develop architecture-aware metrics that can accurately reflect the computational and intelligence trade-offs involved.

OpenAI and other organizations are likely to clarify Astra’s performance through detailed technical disclosures, including real compute costs and latency measurements. Meanwhile, the AI community will scrutinize the validity of token-based efficiency metrics and push for more transparent standards.

Amazon

AI index revision tracking

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why did Astra’s benchmark score change after the index revision?

The score changed because the Artificial Analysis Intelligence Index was updated, which altered the evaluation criteria and scoring benchmarks, making previous scores outdated.

Does Astra perform better or worse than Fable overall?

It depends on the index and task type. Astra shows efficiency gains in coding tasks but is less cost-effective for general intelligence tasks compared to its predecessor, according to AA’s own analysis.

Are token counts a reliable measure of AI compute and efficiency?

No, especially for architectures like Astra that reason in latent space without emitting tokens. Token-based metrics can be misleading for models with complex internal reasoning processes.

Will future benchmarks better reflect Astra’s architecture?

Likely yes. The community is moving toward developing metrics that account for architectural differences, including latent reasoning and non-token-based computation.

What should users and developers take away from this analysis?

They should be cautious about interpreting benchmark scores, especially when models have architectural innovations that traditional metrics do not capture accurately.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
You May Also Like

Create Advanced Chrome Extensions Easily With AI And No-Code Platforms

New no-code platform enables users to create advanced Chrome extensions via AI prompts, removing the need for coding skills and expanding browser automation.

The queue. Why the grid, not the chip, is the binding constraint on AI.

The US interconnection queue now bottlenecks AI infrastructure growth, shifting focus from chips to grid capacity and raising political and economic challenges.

Transform Your Business With Cutting-Edge AI Tools & Automation

Discover how AI tools and automation can optimize workflows, reduce repetitive tasks, and enhance decision-making for your business growth.

AI Adoption: Slow Progress, Lasting Presence

Despite slow adoption, incumbent enterprise platforms like Microsoft and SAP dominate AI integration, highlighting a durable, embedded presence.