3 August 2026 public-source review
3 August 2026 public-source review. Daily public-source review preserved as part of the research history.
Shaduf AI Model Degradation Watch
Cutoff: 2026-08-03 23:59 UTC Scope: Public internet sources only. No Shaduf evaluations, prompts, canaries, benchmarks, or experiments were run.
Current-summary article
The admissible public record does not establish generalized degradation of the underlying model weights across OpenAI, Anthropic, xAI, or the selected open-weight models.
What it does establish is rapid change in the surrounding systems: new model releases, altered pricing, product-specific routing, safety classifiers, fallback models, context/tool changes, and service outages. Those factors can change a user’s observed experience without proving that the base model became less capable.
OpenAI’s GPT-5.6 family became generally available on July 9 as three separately named models: GPT-5.6 Sol, GPT-5.6 Terra, and GPT-5.6 Luna. Their API identifiers are separately documented as gpt-5.6-sol, gpt-5.6-terra, and gpt-5.6-luna. OpenAI explicitly discusses API, ChatGPT, and Codex as distinct surfaces. Its July 30 price-performance update changed Terra and Luna pricing and described improvements spanning models, inference systems, routing, and context management—not only model weights. OpenAI preview, June 26, OpenAI launch, July 9, OpenAI price update, July 30.
OpenAI’s official status record documents a Codex Sol overload incident on July 17, a broader ChatGPT/Codex incident on July 19, and elevated API, ChatGPT, and Codex error rates on July 23–24. These are availability and infrastructure events. OpenAI’s status page itself warns that aggregate availability does not represent every model, subscription, or feature equally. None is evidence that GPT-5.6 Sol, Terra, or Luna’s underlying weights degraded. Codex Sol incident, July 17, July 19 write-up, July 23–24 incident, OpenAI status history.
Anthropic exposes an important counterexample to treating access tiers as independent models. Claude Fable 5 and Claude Mythos 5 share the same underlying model, according to Anthropic, but differ in safeguards and access. Fable adds stronger safety classifiers and may fall back to Claude Opus 4.8 for certain flagged requests; Mythos is a restricted Project Glasswing tier with fewer safeguards. A change in fallback frequency or classifier behavior can therefore look like a model-quality change while the underlying weights remain unchanged. Anthropic launch, June 9, Anthropic safeguards, July 2.
The third monitored Claude row is Claude Opus 5, released July 24 with the API identifier claude-opus-5. Anthropic’s current model documentation positions it as the model for complex agentic coding and enterprise work, below Fable in capability but above Sonnet in the current lineup. Sonnet 5 is current and appears in the August 3 status incident, but it is not included in the three-row Claude monitoring allocation. Anthropic model overview, Anthropic release notes, July 24 entry.
xAI’s current leading model is Grok 4.5. The API identifier is grok-4.5, but the public record does not establish a fixed model ID or snapshot for Grok in X, the consumer Grok website/apps, or the Grok Build coding product. xAI documents Grok 4.5 as the default model in Grok Build, while the consumer and X surfaces remain product-level access paths. This distinction prevents an API result from being silently treated as a result for every Grok surface. xAI launch, July 16, xAI Grok 4.5 API documentation, xAI model catalog.
The three open-weight selections are Kimi K3, GLM-5.2, and DeepSeek V4 Pro. All have public downloadable weights and an identified license. Kimi K3 is open-weight under a custom Kimi K3 license rather than an unqualified OSI open-source license. GLM-5.2 and DeepSeek V4 Pro are released under MIT. The selection is based on current relevance and fresh independent public comparisons available before the cutoff; it is not a claim that these are universally the three best open models. Kimi K3 release record, July 27, GLM-5.2 model card, DeepSeek V4 announcement, April 24.
The strongest independent comparison evidence comes from Artificial Analysis, but its scores are snapshots of complete systems under specified effort, fallback, tool, and harness configurations. A July 24 snapshot placed Opus 5 at 61, Fable 5 at 60, GPT-5.6 Sol at 59, Kimi K3 at 57, and Opus 4.8 at 56 on its Intelligence Index. These figures are comparable within that article, but they are not pure measurements of base weights: Fable and Opus 5 used fallback configurations. Artificial Analysis, July 24.
Likewise, DeepSeek V4 Pro was reported at 52 in an April 24 Artificial Analysis article, while a later public Artificial Analysis page showed 44. Without a frozen methodology and raw historical record, that difference is not proof of degradation. It may reflect endpoint, configuration, benchmark, or index changes. DeepSeek V4 comparison, April 24, current DeepSeek V4 Pro page.
Monitored model comparison
The table contains exactly ten monitored rows. Fable 5 and Mythos 5 are deliberately separate access-tier rows, but they are not claimed to be independent underlying base models.
| # | Monitored model | Explicit surface and identifier | Official release and material changes | Independent public comparison | Official incident / availability record | Cutoff interpretation |
|---|---|---|---|---|---|---|
| 1 | GPT-5.6 Sol | API: gpt-5.6-sol; alias gpt-5.6 in current API documentation. ChatGPT: separate product access. Codex: separate coding surface. | Previewed June 26; GA July 9. July 30 price update retained API pricing at $5 input / $30 output per 1M tokens and introduced Fast mode. Preview, GA launch, API model documentation, price update. | Artificial Analysis: 59 on July 9, max configuration; 59 in the July 24 comparison. July 9 comparison, July 24 comparison. | Codex 5.6-Sol server-overload errors on July 17; broader OpenAI incidents affected multiple surfaces July 19 and July 23–24. July 17 incident, status history. | An availability problem is established; base-weight degradation is not. |
| 2 | GPT-5.6 Terra | API: gpt-5.6-terra. ChatGPT: separate subscription/product access. Codex: separate coding surface. | GA July 9. July 30 API price reduced to $2 input / $12 output per 1M tokens. GA launch, API model documentation, price update. | Artificial Analysis: 55, max configuration, July 9; cost per task $0.55 in the same article. Artificial Analysis, July 9. | No Terra-specific base-quality incident was established. OpenAI’s multi-surface outages remain availability evidence only. | Public evidence supports a separate model and a lower-cost tier, not a measured quality decline. |
| 3 | GPT-5.6 Luna | API: gpt-5.6-luna. ChatGPT: separate subscription/product access. Codex: separate coding surface. | GA July 9. July 30 API price reduced to $0.20 input / $1.20 output per 1M tokens. GA launch, API model documentation, price update. | Artificial Analysis: 51, max configuration, July 9; cost per task $0.21. Artificial Analysis, July 9. | No Luna-specific quality incident was established. | The lower score and lower cost are a dated cross-tier comparison, not evidence of deterioration. |
| 4 | Claude Fable 5 | API: claude-fable-5. Claude.ai, Claude Code, Claude Cowork: GA access, subject to capacity and policy. Cloud: AWS, Google, Microsoft. | GA June 9. Anthropic describes Fable and Mythos as sharing underlying weights. Fable adds stronger safety classifiers and may fall back to Opus 4.8 for flagged cybersecurity, biology, chemistry, or distillation requests. Access was suspended during the June export-control episode and redeployed July 1. Launch, redeployment, safeguards. | Artificial Analysis reported 64.9 at launch on June 9 and 60 in its July 24 snapshot. It separately observed fallback behavior in its testing. June 9 comparison, July 24 comparison. | Fable availability was suspended and restored; the August 3 Anthropic incident concerned Sonnet 5, not Fable. Redeployment, status API. | A changed safeguard or fallback path can alter observed output without a change to the underlying Fable/Mythos weights. |
| 5 | Claude Mythos 5 | API/documented ID: claude-mythos-5; preview/invitation identifier claude-mythos-preview. Surface: Project Glasswing, restricted organizations; not general consumer access. | Released June 9 as a limited Project Glasswing tier. Anthropic describes it as sharing Fable’s underlying model while operating with fewer safeguards. Mythos access was partially restored June 26 after the June suspension. Launch, redeployment, release notes. | Public Artificial Analysis scores generally evaluate the Fable configuration or Fable with fallback, not a clean Mythos-only series. Therefore no independent base-quality trend is established for Mythos. | Anthropic’s July 30 cybersecurity-evaluation disclosure says Mythos 5 was involved in a misconfigured third-party evaluation environment with live Internet access. Production safeguards were not active; this was not a normal production incident. Anthropic disclosure. | Same underlying weights as Fable; different safeguard/access regime. It must not be treated as a separate base model for degradation claims. |
| 6 | Claude Opus 5 | API: claude-opus-5. Claude Platform and cloud partners: AWS, Google, Microsoft. Claude.ai/Code/Cowork: product access subject to plan and capacity. | Released July 24; 1M context and 128K maximum output are documented, with adaptive thinking and effort settings. Anthropic model overview, release notes, July 24. | Artificial Analysis: 61 in the July 24 snapshot, with maximum effort and fallback enabled in the reported configuration. Artificial Analysis, July 24. | Anthropic’s public cybersecurity disclosure refers to Opus 4.7, Mythos 5, and an internal model in the evaluated incidents—not Opus 5 production use. The exact August 3 status incident was Sonnet 5. | A high public comparison score is established for one dated configuration; no longitudinal Opus 5 degradation signal is established. |
| 7 | Grok 4.5 | API: grok-4.5. Grok Build: Grok 4.5 documented as default, but product routing/snapshot is not separately pinned. Grok consumer web/iOS/Android: no public stable model ID established. Grok in X: no public stable model ID established. | API availability documented July 8; public launch July 16. API price $2 input / $6 output per 1M tokens, 500K context, reasoning and tool support. API release notes, API documentation, launch. | Artificial Analysis: 54 Intelligence Index and 76 Coding Agent Index in its July 8 article. The coding result is explicitly a Grok Build surface result. Artificial Analysis, July 8. | xAI’s reviewed status history did not show a July Grok Web incident, but the live page is not a frozen August 3 snapshot. xAI lists Grok Web, iOS, Android, X, and API as distinct services. xAI status, xAI service status. | API, Build, consumer Grok, and X must remain separate evidence surfaces. API quality cannot be generalized to all Grok access paths. |
| 8 | Kimi K3 | Downloadable weights: moonshotai/Kimi-K3. Local surfaces: Transformers, vLLM, SGLang. Hosted API: public API access was documented by the July 17 independent launch report; a frozen provider endpoint ID was not required for the open-weight selection. | Full-weight release record dated July 27. Model card states 2.8T total parameters, 104B active parameters, and 1,048,576 context. Official GitHub, Hugging Face model card, technical report, July 27. | Artificial Analysis: 57 on July 17, before the full-weight release was recorded. Artificial Analysis, July 17. | No model-specific public status history was located in the reviewed official sources. | Genuine open-weight evidence is established. The Kimi K3 license is custom and commercially conditional; it should not be labeled unqualified OSI open source. License. |
| 9 | GLM-5.2 | Downloadable weights: zai-org/GLM-5.2. Local deployment: documented in the model card. Hosted APIs are separate from the local-weight identity. | Released June 16. The official model card identifies MIT licensing, 1M context, and local deployment. Its repository exposes downloadable safetensor shards. Z.ai release page, official model card, weights file tree. | Artificial Analysis: 51 on June 16, described as the leading open-weight result in that dated comparison. Artificial Analysis, June 16. | No model-specific public status history was located in the reviewed official sources. | Downloadable weights and MIT license are established. The current file-tree state is not treated as a historical August 3 snapshot. |
| 10 | DeepSeek V4 Pro | API: deepseek-v4-pro. Downloadable weights: deepseek-ai/DeepSeek-V4-Pro. APP/WEB: separate DeepSeek product surfaces. | V4 Pro and V4 Flash launched April 24 with 1M context and open weights. Pro is 1.6T total / 49B active in the official model card. A July 31 update changed Flash only and stated that Pro API and APP/WEB were unchanged. DeepSeek launch, change log, model card, weights file tree. | Artificial Analysis: 52 in its April 24 launch comparison; a later public page showed 44. The discrepancy is not interpretable as degradation without a frozen method and endpoint record. April 24 comparison, current page. | DeepSeek retired legacy deepseek-chat and deepseek-reasoner on July 24 after prior routing to V4 Flash. That is a routing/API change, not evidence about V4 Pro weights. April 24 notice. | Downloadable MIT-licensed weights are established. Pro and Flash must not be collapsed, and API retirement/routing must not be interpreted as base-model degradation. |
Dated change and incident timeline
| Date | Event | Surface / category | Interpretation | Source |
|---|---|---|---|---|
| 2026-04-24 | DeepSeek V4 Pro and V4 Flash announced with open weights, 1M context, and API identifiers. | Open weights, API | Establishes the DeepSeek V4 model family and separate Pro/Flash surfaces. | DeepSeek V4 announcement |
| 2026-06-09 | Fable 5 general release; Mythos 5 limited Project Glasswing release. | Claude access tiers | Anthropic explicitly states that the two tiers share underlying weights but differ in safeguards and access. | Anthropic launch |
| 2026-06-12 | Anthropic suspended Fable and Mythos access during export-control review. | Availability / policy | Access interruption, not model degradation. | Anthropic redeployment account |
| 2026-06-16 | GLM-5.2 released; MIT-licensed open weights and 1M context documented. | Open weights | Establishes GLM-5.2 as a qualifying open-weight selection. | Z.ai release, model card |
| 2026-06-26 | GPT-5.6 Sol preview announced; Sol, Terra, and Luna described as distinct tiers. | OpenAI model family | Establishes the three requested names before GA. | OpenAI preview |
| 2026-06-26 | Mythos access restored to a set of U.S. organizations. | Claude availability | Demonstrates access-scope change without a base-weight change. | Anthropic redeployment account |
| 2026-06-30 / 2026-07-01 | Export controls lifted; Fable redeployed globally; cloud access restored. | Claude availability | Service and eligibility change. | Anthropic redeployment account |
| 2026-07-08 | Grok 4.5 API availability announced; Artificial Analysis published its independent comparison. | API / comparison | Establishes grok-4.5, its API price, and a dated third-party score. | xAI release notes, Artificial Analysis |
| 2026-07-09 | GPT-5.6 family reached GA; API IDs documented. | OpenAI API, ChatGPT, Codex | Establishes the three separate API model identities and product-surface distinction. | OpenAI launch, API models |
| 2026-07-09 | Artificial Analysis reported Sol 59, Terra 55, Luna 51. | Independent comparison | Dated tier comparison, not longitudinal evidence. | Artificial Analysis |
| 2026-07-16 | xAI publicly described Grok 4.5 as its leading model and made it available in Grok Build and other surfaces. | Grok product access | Confirms Build access, but not a universal fixed consumer/X snapshot. | xAI launch |
| 2026-07-17 | Codex 5.6-Sol experienced server-overload errors. | OpenAI availability | Reliability incident; not evidence of model-quality decline. | OpenAI incident |
| 2026-07-17 | Artificial Analysis reported Kimi K3 at 57 and described the weights as planned for release. | Open-weight comparison | This score predates the July 27 full-weight release record. | Artificial Analysis |
| 2026-07-19 | OpenAI described a ChatGPT/Codex incident caused by regional infrastructure maintenance, database-replica unavailability, and insufficient failover. | OpenAI availability | Infrastructure failure, not model degradation. | OpenAI write-up |
| 2026-07-21 | OpenAI disclosed an internal Hugging Face evaluation security incident involving GPT-5.6 Sol and a prerelease model with reduced cyber refusals. | Evaluation security | A controlled-evaluation safeguard configuration and infrastructure incident; not ordinary production behavior or a quality decline. | OpenAI disclosure |
| 2026-07-23–24 | OpenAI recorded elevated error rates across API, ChatGPT, and Codex. | Multi-surface availability | Service reliability event. | OpenAI incident |
| 2026-07-24 | Opus 5 released with claude-opus-5. | Claude API/cloud/product | Establishes the third monitored Claude model. | Anthropic release notes |
| 2026-07-24 15:59 UTC | DeepSeek retired legacy deepseek-chat and deepseek-reasoner after prior routing to V4 Flash. | API routing | Alias retirement and routing change; not evidence of V4 Pro degradation. | DeepSeek announcement |
| 2026-07-27 16:49 UTC | Kimi K3 technical report and full-weight release record appeared. | Open weights | Establishes cutoff-admissible downloadable-weight evidence. | Kimi K3 technical report |
| 2026-07-30 | OpenAI reduced Terra and Luna API prices; described broader inference/routing/context-management efficiency work. | Pricing / systems | Price and systems change, not a measured base-model quality change. | OpenAI price update |
| 2026-07-30 | Anthropic disclosed three cybersecurity-evaluation incidents, including Mythos 5 in a misconfigured third-party environment. | Evaluation security | Important safeguard/harness evidence; not production degradation. | Anthropic disclosure |
| 2026-07-31 | DeepSeek updated V4 Flash; Pro API and APP/WEB were stated to be unchanged. | API/model routing | Separates Flash changes from Pro. | DeepSeek change log |
| 2026-08-03 15:13:45–15:29:01 UTC | Anthropic recorded “Degraded performance on Claude Sonnet 5”; error rates returned to baseline at 15:20 UTC. | Anthropic availability | A time-bounded Sonnet 5 service incident; Sonnet 5 is not one of the three monitored Claude rows. | Anthropic status API |
| 2026-08-03 | Grok Build changelog listed product version v0.2.120, including model-picker and background-task changes. | Grok Build product | Product UX/task-management change; no fixed underlying model snapshot or base-quality change established. | Grok Build changelog |
Comparable public charts
These charts use only dated public data from comparable sources. They are not Shaduf measurements.
Chart 1 — Artificial Analysis Intelligence Index snapshot
Publication date: July 24, 2026 Scale: 0–65; higher is better. Important configuration note: Fable 5 and Opus 5 include fallback configurations reported by Artificial Analysis.
Claude Opus 5 61 |███████████████████████████████
Claude Fable 5 60 |██████████████████████████████
GPT-5.6 Sol 59 |█████████████████████████████▌
Kimi K3 57 |████████████████████████████▌
Claude Opus 4.8 56 |████████████████████████████
Source: Artificial Analysis, “Opus 5” — July 24, 2026.
Underlying CSV:
model,provider,metric,score,effort_or_configuration,publication_date,source_url
Claude Opus 5,Anthropic,Artificial Analysis Intelligence Index,61,"max; Opus 4.8 fallback enabled",2026-07-24,https://artificialanalysis.ai/articles/opus-5
Claude Fable 5,Anthropic,Artificial Analysis Intelligence Index,60,"max; Opus 4.8 fallback enabled",2026-07-24,https://artificialanalysis.ai/articles/opus-5
GPT-5.6 Sol,OpenAI,Artificial Analysis Intelligence Index,59,"max",2026-07-24,https://artificialanalysis.ai/articles/opus-5
Kimi K3,Moonshot,Artificial Analysis Intelligence Index,57,"as reported by Artificial Analysis",2026-07-24,https://artificialanalysis.ai/articles/opus-5
Claude Opus 4.8,Anthropic,Artificial Analysis Intelligence Index,56,"max",2026-07-24,https://artificialanalysis.ai/articles/opus-5
Chart 2 — GPT-5.6 tier comparison
Publication date: July 9, 2026 Two separate scales: Intelligence Index is higher-is-better; cost per task is lower-is-better.
Artificial Analysis Intelligence Index — higher is better
GPT-5.6 Sol 59 |█████████████████████████████▌
GPT-5.6 Terra 55 |███████████████████████████▌
GPT-5.6 Luna 51 |████████████████████████████▌
Cost per Intelligence Index task, USD — lower is better
GPT-5.6 Sol 1.04 |████████████████████
GPT-5.6 Terra 0.55 |███████████
GPT-5.6 Luna 0.21 |████
Artificial Analysis states that it supported OpenAI’s prerelease evaluation for this launch. These are external evaluator results, not Shaduf tests. Artificial Analysis, July 9, 2026.
Underlying CSV:
model,metric,value,unit,configuration,publication_date,source_url
GPT-5.6 Sol,Artificial Analysis Intelligence Index,59,index,max,2026-07-09,https://artificialanalysis.ai/articles/gpt-5-6-has-landed
GPT-5.6 Sol,Cost per Artificial Analysis Intelligence Index task,1.04,USD,max,2026-07-09,https://artificialanalysis.ai/articles/gpt-5-6-has-landed
GPT-5.6 Terra,Artificial Analysis Intelligence Index,55,index,max,2026-07-09,https://artificialanalysis.ai/articles/gpt-5-6-has-landed
GPT-5.6 Terra,Cost per Artificial Analysis Intelligence Index task,0.55,USD,max,2026-07-09,https://artificialanalysis.ai/articles/gpt-5-6-has-landed
GPT-5.6 Luna,Artificial Analysis Intelligence Index,51,index,max,2026-07-09,https://artificialanalysis.ai/articles/gpt-5-6-has-landed
GPT-5.6 Luna,Cost per Artificial Analysis Intelligence Index task,0.21,USD,max,2026-07-09,https://artificialanalysis.ai/articles/gpt-5-6-has-landed
Chart 3 — Public API list prices at the cutoff
Units: USD per 1M input/output tokens. Scope: API list pricing only; excludes consumer subscriptions, cached-input discounts, provider routing, tool costs, and local open-weight deployment.
Input price — scale: █ ≈ $1
Claude Fable 5 $10.00 |██████████
Claude Mythos 5 $10.00 |██████████
GPT-5.6 Sol $5.00 |█████
Claude Opus 5 $ 5.00 |█████
GPT-5.6 Terra $2.00 |██
Grok 4.5 $2.00 |██
GPT-5.6 Luna $0.20 |▏
Output price — scale: █ ≈ $5
Claude Fable 5 $50.00 |██████████
Claude Mythos 5 $50.00 |██████████
Claude Opus 5 $25.00 |█████
GPT-5.6 Sol $30.00 |██████
GPT-5.6 Terra $12.00 |██▍
Grok 4.5 $6.00 |█▏
GPT-5.6 Luna $1.20 |▏
Sources: OpenAI pricing update, July 30, Anthropic Fable/Mythos launch, June 9, Anthropic Opus 5 release notes, July 24, xAI API release notes, July 8.
Underlying CSV:
surface_model,provider,input_usd_per_1m,output_usd_per_1m,pricing_date,surface,source_url
GPT-5.6 Sol,OpenAI,5,30,2026-07-30,API,https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/
GPT-5.6 Terra,OpenAI,2,12,2026-07-30,API,https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/
GPT-5.6 Luna,OpenAI,0.20,1.20,2026-07-30,API,https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/
Claude Fable 5,Anthropic,10,50,2026-06-09,API,https://www.anthropic.com/news/claude-fable-5-mythos-5?type=company
Claude Mythos 5,Anthropic,10,50,2026-06-09,Project Glasswing API/access tier,https://www.anthropic.com/news/claude-fable-5-mythos-5?type=company
Claude Opus 5,Anthropic,5,25,2026-07-24,API,https://platform.claude.com/docs/en/release-notes/overview
Grok 4.5,xAI,2,6,2026-07-08,API,https://docs.x.ai/developers/release-notes
What is established
- GPT-5.6 Sol, Terra, and Luna are separately named models with separately documented OpenAI API identifiers.
- OpenAI API, ChatGPT, and Codex are distinct surfaces. Service incidents affecting one or more surfaces do not establish base-model degradation.
- Fable 5 and Mythos 5 share underlying weights but intentionally differ in safeguards, fallback behavior, and access.
- Claude Opus 5 is a separate current Claude model with identifier
claude-opus-5. - Grok 4.5 is the current documented leading xAI model, with API identifier
grok-4.5; the API, Grok Build, consumer Grok, and Grok in X are not proven to be identical deployments. - Kimi K3, GLM-5.2, and DeepSeek V4 Pro satisfy the open-weight selection requirement through public downloadable weights and identified licenses.
- GLM-5.2 and DeepSeek V4 Pro are MIT-licensed according to their official model records. Kimi K3 uses a custom license with additional commercial conditions.
- Provider-documented availability, routing, safeguard, and evaluation-harness incidents occurred before the cutoff.
- Artificial Analysis published dated cross-model comparisons before the cutoff.
- Fable’s and Opus 5’s published comparison results include fallback/configuration effects; they are not clean measurements of isolated base weights.
- DeepSeek’s April 24 score and later public score differ, but the evidence does not establish that the difference is degradation.
What is not established
- No public evidence in this run proves generalized degradation across the monitored set.
- No frozen longitudinal dataset compares all ten monitored rows on the same tasks, same surfaces, same settings, and same model snapshots.
- No base-weight quality decline is established for GPT-5.6 Sol, Terra, or Luna.
- No base-weight quality decline is established for Fable/Mythos, Opus 5, or Grok 4.5.
- No public evidence establishes that Grok in X, Grok consumer, and Grok Build use the exact
grok-4.5API snapshot. - No ranking of all ten monitored rows is justified from the available data.
- OpenAI and Anthropic cybersecurity-evaluation disclosures do not establish ordinary production model degradation.
- Outages, elevated error rates, fallback events, or routing changes cannot be converted into quality claims without comparable output evidence.
- Community complaints are not treated as proof because they generally do not identify the exact model snapshot, surface, routing path, system instructions, or fallback state.
Uncertainty
The major uncertainty is not simply statistical noise. It is identity and configuration uncertainty.
A public result may depend on:
- the underlying model snapshot;
- provider routing or alias resolution;
- system instructions;
- safety classifiers and refusal policies;
- fallback models;
- effort or reasoning settings;
- context length and truncation;
- tool availability and tool harness;
- rate limits, overload, or regional capacity;
- API versus consumer-product deployment.
Artificial Analysis’ Fable 5 results illustrate this directly: Fable can fall back to Opus 4.8, so a benchmark result can measure a composite system. The same applies to Opus 5 when fallback is enabled.
Mutable documentation also limits historical reconstruction. Current model pages and model repositories were used only where they expose stable identifiers, licenses, or dated release records. They were not treated as proof that every current field or file-tree state existed unchanged at the cutoff.
The DeepSeek V4 Pro score difference—52 in the April 24 launch comparison versus 44 on a later public page—is a useful warning. Without a frozen methodology, endpoint, configuration, and raw data, it is impossible to distinguish real performance change from measurement drift.
The exact timezone of the displayed timestamps on the OpenAI status pages was not established from the page itself. The Anthropic August 3 incident provides explicit UTC timestamps through its status API.
Freshness
The cutoff rule used here is:
- A claim is admitted when the source was published, updated, or recorded no later than 2026-08-03 23:59 UTC.
- A current page viewed after the cutoff is not used as a historical snapshot unless it exposes a dated release or incident record whose historical state is clear.
- Post-cutoff status entries visible in current provider histories are excluded.
- Current mutable pages are labeled as current or non-frozen where relevant.
- No undocumented model ID, incident, ranking, measurement, or date is inferred.
- All quality numbers are attributed to external provider or evaluator publications; Shaduf did not run tests.
Sources
OpenAI
- GPT-5.6 preview — June 26, 2026
- GPT-5.6 general availability — July 9, 2026
- GPT-5.6 price-performance update — July 30, 2026
- OpenAI API model catalog — current mutable documentation
- GPT-5.6 Sol API page
- GPT-5.6 Terra API page
- GPT-5.6 Luna API page
- OpenAI status history
- Codex Sol overload incident — July 17, 2026
- ChatGPT/Codex infrastructure incident write-up — July 19, 2026
- Elevated error rates — July 23–24, 2026
- Hugging Face model-evaluation security incident — July 21, 2026
Anthropic
- Claude Fable 5 and Mythos 5 — June 9, 2026
- Redeploying Fable 5 — June 30 / July 1, 2026
- Fable safeguards and jailbreak framework — July 2, 2026
- Anthropic model overview — current mutable documentation
- Anthropic release notes
- Anthropic status incident API
- Cybersecurity-evaluation incident disclosure — July 30, updated August 3, 2026
xAI
- Grok 4.5 launch — July 16, 2026
- Grok 4.5 API documentation — updated July 17, 2026
- xAI model catalog
- xAI API release notes
- Grok consumer overview
- Grok consumer FAQ
- xAI/Grok service status
- Grok Web status history
- Grok Build changelog
Open-weight model records
- Kimi K3 official GitHub
- Kimi K3 model card and weights
- Kimi K3 license
- Kimi K3 technical report — July 27, 2026
- GLM-5.2 official release — June 16, 2026
- GLM-5.2 model card
- GLM-5.2 downloadable weights
- DeepSeek V4 launch — April 24, 2026
- DeepSeek API updates
- DeepSeek transparency page
- DeepSeek V4 Pro model card
- DeepSeek V4 Pro downloadable weights
- DeepSeek V4 technical report