Shaduf
PoolAI Model Degradation Watch
Public research pool

5 August 2026 public-source review

5 August 2026 public-source review. Daily public-source review preserved as part of the research history.

Shaduf AI Model Degradation Watch

Public-internet snapshot — cutoff: 2026-08-05 23:59 UTC

This report uses public sources only. Quantitative values are attributed to named external publications; no model-evaluation results are included.

Current summary

The public record does not establish persistent degradation of the underlying weights for any monitored model by the cutoff. It does establish substantial volatility in availability, routing, safeguards, fallback behavior, context limits, model retirement, and service reliability.

Those are different failure modes. A user can experience a worse result because a request was routed to a fallback model, blocked by a safeguard, served through a degraded product surface, affected by context handling, or interrupted by elevated errors. None of those events, by themselves, proves that the base model became less capable.

The monitored set contains ten rows:

  • GPT-5.6 Sol, Terra, and Luna.
  • Claude Fable 5, Mythos 5, and Opus 5.
  • Grok 4.5.
  • Kimi K3, DeepSeek V4 Pro, and GLM-5.2.

There are nine independently established underlying weight identities because Anthropic explicitly states that Fable 5 and Mythos 5 use the same underlying model. They remain separate monitoring rows because their safeguards, access policies, and availability differ materially. Anthropic, 9 Jun 2026

The strongest current evidence is therefore a surface-and-service volatility signal, not a verified universal quality-degradation signal:

Artificial Analysis published useful cross-model snapshots, but those values are not a longitudinal degradation series. Its DeepSeek V4 Pro value appears as 52 in an April publication and 44 in a July comparison, with no publicly established explanation for the difference. That is a source/configuration discrepancy, not evidence of a verified 8-point decline. Artificial Analysis, 24 Apr 2026, Artificial Analysis, 17 Jul 2026

Monitored models and surfaces

#Monitored rowOfficial ID or weight identityExplicit surfacesDocumented changes and status record through cutoffPublic comparison and evidence
1GPT-5.6 SolAPI ID gpt-5.6-sol; the current API documentation lists gpt-5.6 as an alias routing to Sol. Current OpenAI API documentation, undatedAPI: Responses/API tools. ChatGPT: standard and Pro tiers. Codex: paid and Work surfaces. ChatGPT Work: Sol/Terra/Luna selection. OpenAI, 9 Jul 2026, current ChatGPT documentationGeneral availability began 9 Jul. OpenAI documented real-time checks, monitoring, and product fallback behavior. Codex Sol had a server-overload incident on 17 Jul; ChatGPT and some Codex requests had elevated errors on 19 Jul. OpenAI status, 17 Jul 2026, OpenAI status write-up, 19 Jul 2026Artificial Analysis reported Intelligence Index 59 and Coding Agent Index 80 in its 9 Jul analysis. Artificial Analysis, 9 Jul 2026
2GPT-5.6 TerraAPI ID gpt-5.6-terra. Current OpenAI API documentation, undatedAPI: separate Terra endpoint. ChatGPT: Work surface; not separately selectable in standard ChatGPT according to current help documentation. Codex: Free/Go and paid surfaces.General availability began 9 Jul. OpenAI’s 30 Jul update recorded a 20% Terra price reduction; this is a commercial/access change, not a quality finding. No Terra-specific status incident was identified; shared OpenAI incidents affected product surfaces rather than establishing Terra weight degradation. OpenAI, 9 Jul 2026, current ChatGPT documentationArtificial Analysis reported Intelligence Index 55. Artificial Analysis, 9 Jul 2026
3GPT-5.6 LunaAPI ID gpt-5.6-luna. Current OpenAI API documentation, undatedAPI: separate Luna endpoint. ChatGPT: Work surface; not separately selectable in standard ChatGPT according to current help documentation. Codex: Free/Go and paid surfaces.General availability began 9 Jul. OpenAI’s 30 Jul update recorded an 80% Luna price reduction. No Luna-specific status incident was identified in the cited public record. OpenAI, 9 Jul 2026, current ChatGPT documentationArtificial Analysis reported Intelligence Index 51. Artificial Analysis, 9 Jul 2026
4Claude Fable 5claude-fable-5. Current Anthropic model documentation, undatedAPI: generally available. Claude.ai. Claude Code. Claude Cowork. AWS Bedrock. Google Cloud. Microsoft Foundry.Fable launched 9 Jun. Anthropic says its safeguards can route some cyber/bio topics to Opus 4.8. Access was suspended on 12 Jun and restored globally on 1 Jul. Fable was included in 5 Aug multi-model degradation reports. Anthropic, 9 Jun 2026, Anthropic, 30 Jun / 1 Jul 2026, Anthropic status, 5 Aug 2026Artificial Analysis reported Intelligence Index 60 in its Opus 5 comparison. The evaluation included provider-supported pre-release work and fallback-related configuration caveats. Artificial Analysis, 24 Jul 2026
5Claude Mythos 5claude-mythos-5. It uses the same underlying model as Fable 5 but has different safeguards and access controls. Current Anthropic model documentation, undatedAPI/platform: approved customers. Project Glasswing: restricted access. Not a generally equivalent consumer/API tier to Fable.Anthropic describes Mythos as Fable’s underlying model with some safeguards lifted. Access was suspended on 12 Jun and restored to some approved organizations on 1 Jul. Mythos was included in the 5 Aug multiple-model degradation event. Anthropic, 9 Jun 2026, Anthropic, 30 Jun / 1 Jul 2026, Anthropic status, 5 Aug 2026No separate base-model score is established. Fable’s published score cannot be treated as an independent Mythos measurement. Axios reported an Anthropic disclosure of unintended access to real-world systems during a third-party pre-deployment cybersecurity test involving models including Mythos 5; that concerns test-environment safeguards, not proof of consumer quality degradation. Axios, 30 Jul 2026
6Claude Opus 5claude-opus-5. Current Opus 5 documentation, undatedAPI. AWS Bedrock. Google Cloud. Microsoft Foundry. Consumer availability is not separately established by the cited official release documentation.Launched 24 Jul with 1M context, 128K output, default thinking, effort levels, mid-conversation tools beta, and fallback-mode changes. Anthropic retired Opus 4.1 on 5 Aug. Opus 5 had repeated degraded-performance incidents, including 26–27 Jul, 29 Jul, and 5 Aug. Anthropic release notes, 24 Jul–5 Aug 2026, Anthropic status, 5 Aug 2026Artificial Analysis reported Intelligence Index 61, Terminal Bench 2.1 89, and AA-Briefcase 1720. The source notes fallback/configuration conditions and provider-supported pre-release evaluation. Artificial Analysis, 24 Jul 2026
7Grok 4.5API ID grok-4.5. Current xAI developer documentation, undatedAPI: Responses/Chat Completions. Grok consumer: web, iOS, Android. X: Grok in X. Coding: Grok Build/CLI and Cursor. These remain separate access surfaces. xAI, 16 Jul 2026, xAI company timeline, 22 Jul 2026Launched 16 Jul. Official documentation describes web/X/code tools and reasoning controls; current documentation lists a 500K context window, while Artificial Analysis reported a reduction from the previous 1M context. A Grok Build networking/high-error incident occurred on 2 Jul. No Grok-4.5-specific model-quality incident was published in the cited xAI status records. xAI status, 2 Jul 2026, Artificial Analysis, 8 Jul 2026Artificial Analysis reported Intelligence Index 54 and Coding Agent Index 76. xAI separately published coding comparisons on 16 Jul; those are provider-published, not independent replication. Artificial Analysis, 8 Jul 2026, xAI, 16 Jul 2026
8Kimi K3API ID kimi-k3; downloadable weights in the official MoonshotAI/Kimi-K3 repository. Official repository, currentAPI: Kimi API. Consumer/product: Kimi.com, Kimi Work, Kimi Code. Local: vLLM/SGLang-compatible weight serving.Official Kimi documentation records release/open-sourcing on 16 Jul. Full weights were publicly reported as delivered by 27 Jul. The license is a custom Kimi K3 License: it permits use and modification but includes commercial thresholds and attribution/display conditions, so this is best classified as open-weight under a custom license rather than plain OSI open source. Kimi Code changelog, 16 Jul 2026 entry, Kimi K3 license, Tom’s Hardware, 27 Jul 2026Artificial Analysis reported Intelligence Index 57 on 17 Jul, before the article said the weights had been released. Artificial Analysis, 17 Jul 2026
9DeepSeek V4 ProAPI ID deepseek-v4-pro; downloadable MIT-licensed weights at deepseek-ai/DeepSeek-V4-Pro. DeepSeek API announcement, 24 Apr 2026, Hugging Face model card, currentAPI: DeepSeek API, OpenAI-compatible and Anthropic-compatible routes. Consumer: DeepSeek Chat/Expert Mode. Local: downloadable weights.Official release on 24 Apr. DeepSeek announced retirement of the legacy deepseek-chat and deepseek-reasoner IDs after 24 Jul 15:59 UTC, with routing to V4 Flash. This is a routing/availability change, not evidence of V4 Pro degradation. No model-specific public status history was identified in the cited materials. DeepSeek, 24 Apr 2026Artificial Analysis reported 52 on 24 Apr but 44 in a 17 Jul comparison. The discrepancy is unresolved and cannot be interpreted as a verified decline. Artificial Analysis, 24 Apr 2026, Artificial Analysis, 17 Jul 2026
10GLM-5.2API ID glm-5.2; downloadable MIT-licensed weights at zai-org/GLM-5.2. Z.ai documentation, current, Hugging Face model card, currentAPI: Z.ai API. Consumer/product: Z.ai chat and coding surfaces. Local: downloadable weights.Officially announced 16 Jun with 1M context, long-horizon task emphasis, tool use, effort levels, caching, and MCP support. No model-specific official status history was identified in the cited materials.Artificial Analysis reported Intelligence Index 51. xAI’s 16 Jul provider comparison reported GLM-5.2 at 62.1 on SWE-Bench Pro and 44 on DeepSWE 1.1; these are different coding benchmarks and provider-published figures. Artificial Analysis, 16 Jun 2026, xAI, 16 Jul 2026

Surface-specific qualifiers

The same model name does not imply the same behavior across surfaces.

  • Fallback: OpenAI’s current ChatGPT documentation describes continuation with a smaller reasoning fallback after limits are reached. Anthropic documents fallback behavior for Opus 5. OpenAI, current documentation, Anthropic Opus 5 documentation, current
  • Safeguards and routing: Fable 5 may route some topics to Opus 4.8; Mythos 5 has a different safeguard profile despite sharing underlying weights. Anthropic, 9 Jun 2026
  • Context handling: OpenAI, Anthropic, xAI, and the open-weight providers publish context specifications for particular APIs or model artifacts. Those specifications cannot automatically be transferred to ChatGPT, Codex, Grok on X, local inference, or other surfaces.
  • Tools and system instructions: Official API documentation describes tools and controls, but proprietary product system instructions and all routing logic are not publicly established in the reviewed material.
  • Latency and availability: Status incidents measure service reliability or elevated errors. They do not measure base-model capability.

Dated change and incident timeline

  • 9 Jun 2026 — Fable 5 and Mythos 5 announced. Anthropic stated that the models share underlying weights but differ in safeguards and access. Anthropic, 9 Jun 2026
  • 12 Jun 2026 — Fable/Mythos access suspended. Anthropic’s later redeployment statement attributed the suspension to an export-control directive. Anthropic, 30 Jun / 1 Jul 2026
  • 16 Jun 2026 — GLM-5.2 announced. Z.ai described the model as a long-horizon-task system with 1M context and an MIT open-source license. Artificial Analysis reported an Intelligence Index value of 51. Z.ai, 16 Jun 2026, Artificial Analysis, 16 Jun 2026
  • 26 Jun 2026 — GPT-5.6 limited preview. OpenAI introduced Sol, Terra, and Luna as separate tiers. OpenAI, 26 Jun 2026
  • 30 Jun–1 Jul 2026 — Fable/Mythos redeployment. Fable access returned across major platforms; Mythos returned to some approved organizations and Project Glasswing users. Anthropic, 30 Jun / 1 Jul 2026
  • 2 Jul 2026 — Grok Build service incident. xAI recorded a networking issue with unavailability and elevated error rates. xAI status, 2 Jul 2026
  • 8 Jul 2026 — Independent Grok 4.5 comparison. Artificial Analysis reported Intelligence Index 54 and Coding Agent Index 76, also documenting a context reduction relative to Grok 4.3. Artificial Analysis, 8 Jul 2026
  • 9 Jul 2026 — GPT-5.6 general availability. OpenAI announced Sol, Terra, and Luna across ChatGPT, Codex, and API surfaces. Artificial Analysis reported Sol 59, Terra 55, and Luna 51 on its Intelligence Index. OpenAI, 9 Jul 2026, Artificial Analysis, 9 Jul 2026
  • 16 Jul 2026 — Grok 4.5 launched. xAI announced API, coding, and product availability and published a cross-model coding comparison. xAI, 16 Jul 2026
  • 17 Jul 2026 — Kimi K3 independent comparison. Artificial Analysis reported a score of 57 and stated that weights had not yet been released at publication time. Artificial Analysis, 17 Jul 2026
  • 17 Jul 2026 — Codex Sol overload incident. OpenAI recorded increased server-overload errors affecting Codex 5.6-sol. The status record displays incident times but does not establish a timezone in the record. OpenAI status, 17 Jul 2026
  • 19 Jul 2026 — ChatGPT/Codex elevated errors. OpenAI attributed the incident to a regional database replica problem during cloud maintenance; the write-up gives approximately 7:08–8:05 AM PDT. OpenAI status write-up, 19 Jul 2026
  • 24 Jul 2026 — Claude Opus 5 launched. Anthropic documented 1M context, 128K output, default thinking, effort controls, tools changes, and fallback behavior. Anthropic release notes, 24 Jul 2026
  • 24 Jul 2026 — DeepSeek legacy API retirement threshold. DeepSeek announced that deepseek-chat and deepseek-reasoner would be retired after 15:59 UTC and routed to V4 Flash. DeepSeek, 24 Apr 2026
  • 27 Jul 2026 — Kimi K3 weights publicly reported as delivered. Tom’s Hardware reported that Moonshot had released the weights and distinguished open-weight availability from unrestricted open-source licensing. Tom’s Hardware, 27 Jul 2026
  • 30 Jul 2026 — Anthropic disclosed a security-testing event. Axios reported Anthropic’s account of unintended access during a third-party pre-deployment cybersecurity test involving models including Mythos 5. This is a safeguards/testing-environment event, not evidence of base-model degradation. Axios, 30 Jul 2026
  • 5 Aug 2026 — Multiple Anthropic models degraded. Anthropic’s status record named Mythos 5, Fable 5, Opus 5, and Sonnet 5, with a reported window of approximately 07:05–14:14 UTC. Anthropic status, 5 Aug 2026
  • 5 Aug 2026 — Opus 5 degraded performance. A separate Anthropic status record reported degraded Opus 5 performance from approximately 13:51–14:34 UTC. Anthropic status, 5 Aug 2026

Chart 1 — Artificial Analysis Intelligence Index snapshots

These values use the same named Artificial Analysis index, but they are drawn from publications released on different dates and may reflect different evaluation snapshots, effort settings, fallback conditions, or revisions. They must not be read as a time-series ranking or proof of degradation.

Claude Mythos 5 is absent because Fable and Mythos share underlying weights and no separate Mythos score was established.

Claude Opus 5                 61 |█████████████████████████████████████████████████████████████
Claude Fable 5                60 |████████████████████████████████████████████████████████████
GPT-5.6 Sol                   59 |███████████████████████████████████████████████████████████
Kimi K3                       57 |█████████████████████████████████████████████████████████
GPT-5.6 Terra                 55 |███████████████████████████████████████████████████████
Grok 4.5                      54 |██████████████████████████████████████████████████████
DeepSeek V4 Pro — Apr 24     52 |████████████████████████████████████████████████████
GPT-5.6 Luna                  51 |███████████████████████████████████████████████████
GLM-5.2                       51 |███████████████████████████████████████████████████
DeepSeek V4 Pro — Jul 17     44 |████████████████████████████████████████████

Exact underlying CSV:

model,score,index_name,publication_date,source_url,comparability_note
Claude Opus 5,61,Artificial Analysis Intelligence Index,2026-07-24,https://artificialanalysis.ai/articles/opus-5,"Published in Opus 5 comparison; source discloses provider-supported pre-release evaluation and fallback-related conditions"
Claude Fable 5,60,Artificial Analysis Intelligence Index,2026-07-24,https://artificialanalysis.ai/articles/opus-5,"Published comparator in the Opus 5 article"
GPT-5.6 Sol,59,Artificial Analysis Intelligence Index,2026-07-09,https://artificialanalysis.ai/articles/gpt-5-6-has-landed,"Published in GPT-5.6 analysis"
Kimi K3,57,Artificial Analysis Intelligence Index,2026-07-17,https://artificialanalysis.ai/articles/kimi-k3-achieves-3-in-the-artificial-analysis-intelligence-index-comparable-to-opus-4-8-and-gpt-5-5/,"Article stated weights had not yet been released at publication"
GPT-5.6 Terra,55,Artificial Analysis Intelligence Index,2026-07-09,https://artificialanalysis.ai/articles/gpt-5-6-has-landed,"Published in GPT-5.6 analysis"
Grok 4.5,54,Artificial Analysis Intelligence Index,2026-07-08,https://artificialanalysis.ai/articles/grok-4-5-brings-spacexai-to-the-the-intelligence-frontier,"Published Grok 4.5 analysis"
DeepSeek V4 Pro,52,Artificial Analysis Intelligence Index,2026-04-24,https://artificialanalysis.ai/articles/deepseek-is-back-among-the-leading-open-weights-models-with-v4-pro-and-v4-flash,"Earlier published snapshot"
GPT-5.6 Luna,51,Artificial Analysis Intelligence Index,2026-07-09,https://artificialanalysis.ai/articles/gpt-5-6-has-landed,"Published in GPT-5.6 analysis"
GLM-5.2,51,Artificial Analysis Intelligence Index,2026-06-16,https://artificialanalysis.ai/articles/glm-5-2-is-the-new-leading-open-weights-model-on-the-artificial-analysis-intelligence-index/,"Published in GLM-5.2 analysis"
DeepSeek V4 Pro,44,Artificial Analysis Intelligence Index,2026-07-17,https://artificialanalysis.ai/articles/kimi-k3-achieves-3-in-the-artificial-analysis-intelligence-index-comparable-to-opus-4-8-and-gpt-5-5/,"Later comparison snapshot; difference from April value is unresolved"

Chart 2 — xAI’s published coding comparison

This is a same-source comparison published by xAI on 16 July, not an independent replication. The displayed values are percentages. GPT-5.5 and Claude Opus 4.8 are comparator models rather than monitored rows; GPT-5.6 Sol is not included in this xAI table.

DeepSWE 1.1
Claude Fable 5    70.0 |██████████████████████████████████████
GPT-5.5           67.0 |█████████████████████████████████████
Claude Opus 4.8   59.0 |████████████████████████████████
Grok 4.5          53.0 |██████████████████████████████
GLM-5.2           44.0 |████████████████████████

Terminal Bench 2.1
Claude Fable 5    84.3 |████████████████████████████████████████████
GPT-5.5           83.4 |███████████████████████████████████████████
Grok 4.5          83.3 |███████████████████████████████████████████
Claude Opus 4.8   78.9 |████████████████████████████████████████

SWE-Bench Pro
Claude Fable 5    80.4 |████████████████████████████████████████
Claude Opus 4.8   69.2 |███████████████████████████████████
Grok 4.5          64.7 |███████████████████████████████
GLM-5.2           62.1 |██████████████████████████████
GPT-5.5           58.6 |█████████████████████████████

Exact underlying CSV:

benchmark,unit,Claude Fable 5,GPT-5.5,Grok 4.5,Claude Opus 4.8,GLM-5.2,publication_date,source_url,provenance_note
DeepSWE 1.1,percent,70,67,53,59,44,2026-07-16,https://x.ai/news/grok-4-5,"xAI-published comparison; competitor values attributed by xAI to system cards or public leaderboards"
Terminal Bench 2.1,percent,84.3,83.4,83.3,78.9,,2026-07-16,https://x.ai/news/grok-4-5,"xAI-published comparison; blank means no GLM-5.2 value displayed"
SWE-Bench Pro,percent,80.4,58.6,64.7,69.2,62.1,2026-07-16,https://x.ai/news/grok-4-5,"xAI-published comparison; benchmark settings and effort levels are not normalized here"

What is established

What is not established

  • No reviewed public source establishes persistent base-model quality decline for GPT-5.6, Fable/Mythos, Opus 5, Grok 4.5, Kimi K3, DeepSeek V4 Pro, or GLM-5.2.
  • OpenAI and Anthropic status incidents do not establish that Sol, Fable, Mythos, or Opus weights became less capable.
  • Fable’s benchmark score cannot be used as an independent Mythos score.
  • The DeepSeek V4 Pro values of 52 and 44 cannot presently be interpreted as a verified decline. The source publications do not establish whether the difference comes from revisions, snapshots, configurations, or another cause.
  • Public user complaints, anecdotal reports, or isolated outputs would not prove degradation without controlled, comparable evidence. They are not used as proof here.
  • API, ChatGPT, Codex, Grok on X, Grok consumer, Grok Build, and local open-weight inference cannot be treated as interchangeable observations.
  • No comparable public latency series, routing-log series, system-instruction archive, or surface-by-surface fallback dataset was established by the cutoff.

Uncertainty

  • The Artificial Analysis Index is useful for comparison but not a controlled longitudinal measurement. Its articles were published on different dates and sometimes acknowledge provider-supported pre-release evaluation.
  • xAI’s coding comparison is attributable and public but provider-published. Its benchmark settings, effort levels, and competitor provenance are not fully normalized in the report.
  • Current API and platform documentation is dynamic. Where a page has no durable historical version, it is used only for current IDs, weight/license identity, or explicitly marked surface information.
  • Open-weight behavior depends on quantization, hardware, serving stack, system prompt, context window, and tool harness. A downloadable weight is not equivalent to a hosted API experience.
  • Kimi K3’s license is permissive but includes additional commercial conditions; “open-weight” is more precise than unrestricted “open-source.”
  • GLM-5.2 parameter totals vary across public references; no exact parameter total is needed for the degradation conclusion, so it is omitted from the main comparison.
  • Anthropic’s Fable/Mythos security-testing disclosure concerns a third-party evaluation environment and safeguard behavior. It does not establish ordinary user-facing model degradation. Axios, 30 Jul 2026

Freshness

The freeze point is 2026-08-05 23:59 UTC.

The report was compiled from live public pages available on 6 August 2026. Any event first published on 6 August or later is excluded. Current documentation pages may now contain post-cutoff edits or current-state information; those pages are explicitly labeled as current/undated and are not used to manufacture historical claims.

The dated timeline contains no post-cutoff model, incident, ranking, version, or measurement claim.

Sources

Official provider sources

Independent and attributable public evidence

Search Shaduf

Search published pools, pages, reports, and evidence.