Back to today's topics

Verified · Aug 5, 2026

Independently verified

Ring-Zero 7/16 (arXiv 2607.12395): 1T-parameter zero RL reaches 94.1% on AIME 2024, 5 spontaneously emergent cognitive behaviors

2 sources

arXiv 2607.12395 (indexed as a Hugging Face daily paper on 2026-07-16) introduces Ring-2.5-1T-Zero, a model trained with zero reinforcement learning applied directly to a pretrained base model without human-annotated CoT data, relying on RL with verifiable rewards (RLVR). Architecture: two pretrained base models from the Ling-2.5 family — Ling-2.5-1T-Base (1T MoE, 63B activated) and Ling-2.5-flash-Base (104B MoE, 7.4B activated). Four-stage training pipeline: (1) first-stage RL with clipped importance sampling policy gradient and token-level loss to elicit reasoning; (2) self-distillation compressing verbose CoT traces via shortest-correct-rollouts + self-refinement; (3) second-stage RL with sample-level loss normalization for stable, sustained optimization; (4) third-stage RL with tier-based training at three difficulty tiers (4k / 16k / 64k token budgets). Infrastructure: 320x H200 GPUs with Megatron as the training engine, SGLang for rollout, Areal framework for orchestration. Optimizations include mixed-precision control (BF16 body with FP32 attention softmax and LM head) and context parallelism tailored to the model's hybrid MLA + Lightning Attention architecture. Findings: validates the 'bitter lesson' at 1T-parameter scale; reveals a two-phase training dynamic (initial discovery + later sharpening); Ring-2.5-1T-Zero (Second Stage RL) reaches 94.1% on AIME 2024 with competitive numbers on AIME 2025-2026 / HMMT 2025-2026 / IMOAnswerBench; spontaneously develops five cognitive behaviors without explicit supervision: anthropomorphism, structured formatting, self-verification, parallel reasoning, and context anxiety; CoT traces outperform baselines on comprehensibility, reproducibility (5.8-point gain on Qwen-32B distillation vs DeepSeek-R1), and efficiency (less than half the tokens of competitors).

Why now

Ring-Zero is the strongest zero RL + 1T-parameter research of early July — without human-annotated CoT data, a 1T-parameter model hits 94.1% on AIME 2024 and spontaneously develops 5 cognitive behaviors. Creators can frame this as 'in the RLVR era, 1T-parameter zero RL also spontaneously develops cognitive behaviors.'

Why it is worth publishing

Big demo surface: pull Ring-2.5-1T-Zero's 5 emergent cognitive behaviors (anthropomorphism / structured formatting / self-verification / parallel reasoning / context anxiety) into a comparison card, and explain 'why zero RL training can spontaneously develop behaviors that don't have explicit supervision.'

Evidence basis

arXiv 7/16 indexing + 1T-parameter + 94.1% on AIME 2024 + 5 emergent cognitive behaviors — heat is medium-to-high as a single research event.

Ring-Zero just hit arXiv — 1T-parameter zero RL reached 94.1% on AIME 2024, and spontaneously developed 5 cognitive behaviors (anthropomorphism / self-verification / parallel reasoning / etc.) without any human-annotated CoT data.

Angle

Frame Ring-Zero (arXiv 2607.12395) as 'the strongest zero RL + 1T-parameter research of early July' — bundle 'in the RLVR era, 1T-parameter zero RL also spontaneously develops cognitive behaviors' into one piece rather than reading AIME 2024 94.1% in isolation.

Format

Long-form explainer

Demo idea

Record a 10-minute four-segment demo: 3 minutes on 'why zero RL scales to 1T' (Ling-2.5-1T-Base / Ling-2.5-flash-Base + RLVR without human-annotated CoT); 3 minutes on the four-stage training pipeline (first-stage RL with clipped importance sampling / self-distillation compressing CoT / second-stage RL with sample-level loss normalization / third-stage RL with three 4k/16k/64k token-budget tiers); 2 minutes on the 5 emergent cognitive behaviors (anthropomorphism / structured formatting / self-verification / parallel reasoning / context anxiety) and 'why zero RL can spontaneously develop them'; 2 minutes on the CoT-trace edge over baselines (comprehensibility / 5.8-point reproducibility gain on Qwen-32B vs DeepSeek-R1 / less than half the tokens).

Platform notes

The paper's AIME 2024 94.1% and the 5 emergent cognitive behaviors are paper self-reported (medium risk) — don't paraphrase as a peer-reviewed or industry-standard taxonomy; specific training compute and reward-formulation details are not in the captured summary (medium risk) — don't fill in from memory.

Usable claims

  • arXiv 2607.12395 (indexed as a Hugging Face daily paper on 2026-07-16) introduces Ring-2.5-1T-Zero, a model trained with zero reinforcement learning applied directly to a pretrained base model without human-annotated CoT data, relying on RL with verifiable rewards (RLVR); the architecture uses two pretrained base models from the Ling-2.5 family — Ling-2.5-1T-Base (1T MoE, 63B activated) and Ling-2.5-flash-Base (104B MoE, 7.4B activated) — and a four-stage training pipeline of (1) first-stage RL with clipped importance sampling policy gradient and token-level loss, (2) self-distillation compressing verbose CoT traces via shortest-correct-rollouts + self-refinement, (3) second-stage RL with sample-level loss normalization, and (4) third-stage RL with tier-based training at three difficulty tiers (4k / 16k / 64k token budgets); infrastructure uses 320x H200 GPUs with Megatron as the training engine, SGLang for rollout, and the Areal framework for orchestration; mixed-precision BF16 body with FP32 attention softmax and LM head; context parallelism tailored to the model's hybrid MLA + Lightning Attention architecture; key findings include validation of the 'bitter lesson' at 1T-parameter scale, a two-phase training dynamic (initial discovery + later sharpening), Ring-2.5-1T-Zero (Second Stage RL) reaching 94.1% on AIME 2024 with competitive numbers on AIME 2025-2026 / HMMT 2025-2026 / IMOAnswerBench, and the spontaneous emergence of five cognitive behaviors without explicit supervision: anthropomorphism, structured formatting, self-verification, parallel reasoning, and context anxiety; CoT traces outperform baselines on comprehensibility, reproducibility (5.8-point gain on Qwen-32B distillation vs DeepSeek-R1), and efficiency (less than half the tokens of competitors).

Evidence pipeline

Breakdown

Reading any single Ring-Zero fact point (1T MoE / 63B activated / 4-stage training pipeline / AIME 2024 94.1% / 5 emergent cognitive behaviors / Qwen-32B distillation 5.8-point gain over DeepSeek-R1 / less than half the tokens) in isolation turns into 'single-benchmark reading.' This piece explains how to use the 'in the RLVR era, 1T-parameter zero RL also spontaneously develops cognitive behaviors' frame — bundle 'why zero RL training can spontaneously develop behaviors that don't have explicit supervision' into one piece; thread 4-stage training pipeline + 5 emergent cognitive behaviors + CoT-trace edge over baselines into a single narrative.

Risks

  • Pin links to each source; quote only what the captured summary states; do not paraphrase specific benchmark numbers, performance metrics, paper claims, integration milestones, or architectural details beyond what is stated.
  • Pin the link to the arXiv paper page; quote only what the paper summary states; do not paraphrase the emergent-behavior taxonomy as a peer-reviewed or industry-standard framework; flag that the headline scores are paper self-reported.

Demo ideas

  • Build a 'Ring-Zero 5 emergent cognitive behaviors' information card: pair each behavior with a concrete demo prompt (anthropomorphism / structured formatting / self-verification / parallel reasoning / context anxiety).
  • Record a 'why zero RL scales to 1T' demo — walk through the Ling-2.5-1T-Base + SGLang + Areal training-stack details and explain what 'no human-annotated CoT data' means in the RLVR era.
  • Build a four-stage training-pipeline flowchart: first-stage RL with clipped importance sampling → self-distillation compressing CoT → second-stage RL with sample-level loss normalization → third-stage RL with three 4k/16k/64k token-budget tiers.