Tev: Scaling Long-Horizon Agentic Environments with Dense Progress Rewards

1Adobe Research 2UCLA University of California, Los Angeles
Paper arXiv soon Code soon Models & Data soon X soon BibTeX
Left: Tev versus Qwen base on ALE-CM, TB 3.0-CM, TB 4.0-CM and TB 2.1, and all models on ALE-CM ranked by Pass@5. Right: the latent-world-first pipeline, from latent world to task environment to progress reward.
Overview. Left: Tev's results on ALE-CM and its zero-shot generalization to Terminal-Bench (TB) 3.0-CM, 4.0-CM and 2.1, across two model scales. Tev-27B matches GPT-6 Astra and Claude Opus 5 at Pass@5 on ALE-CM. Right: our latent-world-first data synthesis, in brief.
TL;DRWe build long-horizon agent tasks backward from a hidden, known-correct latent world, so every task comes with an exact verifier and a dense, per-step progress reward. RL with that reward lifts open 9B, 27B and 320B models on long-horizon benchmarks, including ones never seen in training.
1,170verified environmentsacross 7 domains, the first long-horizon training set with progressive rewards
12.4 minmedian agent rolloutabout 6× longer than the longest prior corpus (SWE-Smith, 2.0 min)
0.28ALE-CM Pass@5Tev-27B, +65% over Qwen3.8-27B; matches GPT-6 Astra and Claude Opus 5
2.5×TB 4.0-CM Pass@10.04 → 0.10 with no Terminal-Bench task used for training
Abstract

Abstract

Long-horizon agents have drawn growing attention for their ability to work autonomously—from hours to days—on economically valuable tasks, a capability essential to automated discovery and recursive self-improvement (RSI). Yet open-source models still struggle on long-horizon tasks, largely for lack of both training data and a working recipe, and scaling both is hard: long, complex tasks are difficult to curate and verify, and their rewards are sparse. We address both. First, we introduce Tev-Gym, the first long-horizon agent training dataset with verifiable, dense, progressive rewards—1,170 verified task environments across 7 domains. It is built on a latent-world-first approach that first constructs a latent world: a complete and correct hidden specification of the task domain, workflow, and expected state after each step. Each task is then progressively rendered from this specification into an agent-facing environment, enabling controllable complexity and a natural progress metric.

Second, we turn this metric into a training recipe: via progress-potential reward shaping (PPRS), we convert progress into a dense per-step reward that improves training efficiency and makes otherwise-intractable tasks trainable. Training with this recipe yields Tev-27B, which achieves substantial gains on the Agents' Last Exam computing and math sciences domains (ALE-CM), matching GPT-6 Astra and Claude Opus 5 at Pass@5. Despite training only on ALE-related data, it generalizes to the Terminal-Bench (TB) 3.0 and 4.0 subsets, surpassing GLM 5.2 and Gemini 3.7 Flash, and reaches Claude Opus 4.8-level performance on middle-horizon tasks such as TB 2.1.

Video

The paper in under three minutes

A walkthrough of the paper: the two obstacles to long-horizon training, latent-world-first synthesis, the progress reward, the results and the ablations.
Motivation

Why long-horizon agent training is hard

1No long-horizon training corpus

Open agentic corpora such as SWE-Smith, TMAX, Endless Terminals, CLI-Gym, SETA and Terminal-Universe hold thousands of environments, but each is a pull-request-sized fix or a short terminal session, graded by unit tests with a pass/fail signal. Genuinely long-horizon tasks exist only in small, human-authored benchmarks that take weeks of expert effort per task and are meant for evaluation, not training.

2Terminal reward starves RL

GRPO normalizes rewards within a group of \(G\) rollouts of the same task. Under a terminal-only reward, a group in which every rollout fails has zero advantage everywhere and contributes nothing to the gradient. If the policy solves a task with probability \(p\), a fraction \((1-p)^G\) of groups is wasted, and on long-horizon tasks an untrained open model solves roughly one in ten or fewer.

How many groups carry no signal?

Each square is one group of \(G\) rollouts. Gray groups have no successful rollout, so terminal-only reward gives every rollout in them zero advantage.

66%of groups are wasted under terminal-only reward, \((1-p)^G\). The paper's examples: 66% at \(p=0.05, G=8\) and 92% at \(p=0.01\). With Tev's progress reward, a group still gets signal whenever its rollouts progress to different depths.
no rollout passes: zero gradientat least one pass

Tev-Gym versus prior agentic task corpora

DatasetSize# DomainsAgent time (min)OracleGrading metricProgressive
reward
P@1P@8
MedianMean
SWE-Smith50.1k12.02.4unit testspass/fail✗47%52%
Terminal-Universe32.0k3NANApytestpass/fail✗NANA
TMAX-15K14.6k31.32.1unit testspass/fail✗74%80%
SETA4.5k41.22.1unit teststest fraction✗80%88%
Endless Terminals3.3k20.50.7unit testspass/fail✗83%89%
CLI-Gym1.7k41.92.3unit testspass/fail✗69%74%
Tev-Gym (ours)1,170712.414.8latent worldstaged Φ + pass/fail✓29%40%

Agent time is the median and mean wall-clock minutes per rollout of Gemini 3.6 Flash, from first action to last step. Difficulty is Pass@k of Gemini 3.6 Flash on every dataset. Tev-Gym tasks run the longest and are the hardest to solve.

Method · Data

Latent-world-first synthesis

Pipelines that ask a language model to write the task, its inputs and its checker have no true answer to check against. We build the answer first: a latent world \(W\), the complete and correct hidden state of the task domain. The environment is then rendered from \(W\) and corrupted in declared, invertible ways, so \(W\) is an exact answer key for the final output and for every intermediate state.

Pipeline: a coding agent writes a seeded generator that builds the latent world (oracle plus corruption plan), compiles a workflow DAG and progress contract, renders clean inputs, applies a corruption ledger, and exposes the public task while graders stay sealed.
Latent-world-first synthesis, on one Tev-Gym environment. The task: “Retail Warehouse ETL Recovery. Goal: build a clean star-schema SQLite warehouse from six months of messy retail data and produce two truthful JSON sidecars.”
  1. Task sourcesTraining data targets ALE-CM, grouped into seven domains. We sample specifications within each domain and author a generator around each one.
  2. World constructionA seeded program builds the latent world \(W\) for a given seed and scale tier (e.g. a warehouse's customers, products and transactions), checks its invariants, and draws a defect plan.
  3. Workflow graphThe workflow is compiled by code, with no model in the loop, into a DAG of dependent stages. Its critical path sets the task's horizon, and its stages form the skeleton of the progress reward.
  4. Artifact rendering\(W\) is projected into the task's input files: CSV, JSON, SQLite, safetensors shards, stripped binaries or infrastructure manifests. Rendering must round-trip back to \(W\) exactly.
  5. Controlled corruptionThe defect plan injects schema drift, duplicate and missing records, category drift, scrambled shard indices, stale caches and more. The corrupted tree \(s_0\) is the task; the untouched \(W\) is the ground truth.
  6. Public spec & hidden suiteThe prompt \(x\) is the only description the agent sees. A sealed suite of metamorphic transforms and plausible-but-wrong counterfactuals is used at verification time.
Every component of an environment is a function of the same latent world: $$E \;=\; \bigl(x,\; s_0,\; \mathrm{ver}_W,\; \Phi_W\bigr)$$ the task prompt \(x\), the corrupted input tree \(s_0\), the strict pass/fail verifier \(\mathrm{ver}_W\) and the progress potential \(\Phi_W\). This single anchor prevents artifact drift, where a separately generated environment, task and verifier each look valid on their own yet describe different problems.
DiverseEvery environment is labeled on 19 coverage axes: subdomain, scale, four horizon axes, workflow stage, tool capability, artifact role, relation type, corruption class and more.
Controllable difficultyScale tiers (unit, small, full) grow \(W\) and the operational horizon; the DAG and the number of corruption classes set the dependency depth.
QualifiedAn environment is admitted only if it passes all 8 automatic qualification gates, including 14 gaming probes that must all score 0.

Seven domains:

  • Software engineering
  • Data engineering
  • AI/CS research
  • Mathematics & operations research
  • Security & forensics
  • Infrastructure administration
  • Quantum computing

The pipeline is not tied to ALE-CM. The paper's appendix walks through three more environments: AWS cost optimization with a billing-dashboard screenshot rendered from the latent world, closed-loop EnergyPlus building control, and financial-statement reconstruction from SEC-style filings.

Method · Reward

Progress-potential reward shaping (PPRS)

The latent world that grades task completion also grades its middle. Every Tev-Gym environment ships a progress contract: nodes \(v \in V\), each a checkable state of the output workspace taken from the workflow DAG, with weights \(w_v\) summing to 1 and prerequisite sets. A partial scorer \(c_v(s) \in [0,1]\) grades the agent's current workspace \(s\) against \(W\), and credit is prerequisite-gated: \(\tilde c_v(s)\) is the lowest partial score among \(v\) and all of its prerequisites.

Progress potential $$\Phi(s) \;=\; \sum_{v \in V} w_v\, \tilde c_v(s) \;\in\; [0,1]$$

Rises as the workflow is completed and falls on regressions. Because \(W\) exists before the task, every correct intermediate state is known exactly: no human step labels, trained reward model or LLM judge.

Shaped rollout reward $$\begin{aligned} R(\tau) &= \mathbb{1}[\tau \text{ passes}] + \lambda \sum_{t=1}^{T} \bigl(\Phi(s_t) - \Phi(s_{t-1})\bigr) \\ &= \mathbb{1}[\tau \text{ passes}] + \lambda\, \Phi(s_T) \end{aligned}$$

With \(\lambda = 0.25\). Since \(\Phi \le 1\), any passing rollout earns \(1+\lambda\), more than any failing one. A group where nothing passes still has rewards \(\lambda\Phi(s_T^{(i)})\) that rank rollouts by how far they got.

Step plot of the progress potential over 30 turns for eight rollouts: terminal-only reward is zero for all eight, while the progress potential separates them into levels 0, 0.26, 0.69 and 1.
The progress potential along one group (\(G=8\)) of rollouts. Eight rollouts of the trained 27B model on one warehouse environment, graded live after every turn. Under terminal-only reward (red) every rollout scores 0. Under progressive rewards, two submit an incomplete warehouse (circles), five exhaust the 65k-token context (crosses), and one builds a complete warehouse, so the rewards vary within the group.
AnchoredThe untouched input tree and every gaming probe score \(\Phi = 0\), and \(\Phi = 1\) exactly when the strict verifier accepts the workspace.
Monotone\(\Phi\) lies strictly between 0 and 1 on partially correct states and rises monotonically to 1 along the reference trajectory.
IndependentThe progress scorer, strict verifier and reference solvers have disjoint code closures, and none of the grading code ships in the public tree.
Experiments

Results

We train Qwen3.5-9B on the 264-task unit tier and Qwen3.8-27B on the full Tev-Gym corpus with DPPO and the shaped reward above: 24 rollouts for each of 8 tasks per step, a 65,536-token budget per rollout, at most 64 tool calls, and a single bash tool. To test whether the recipe transfers to a larger base model, we also train GLM-5.3-Flash the same way, giving Tev-320B, which we evaluate on the TB 3.0 and 4.0 subsets. We evaluate in-domain on ALE-CM, the computing-and-math split of Agents' Last Exam, and zero-shot on the Terminal-Bench 3.0 and 4.0 “CM” subsets (22 and 18 tasks) and the mid-horizon TB 2.1. The subsets track full-set performance (Pearson \(r = 0.96\) and \(0.97\) for TB 3.0 and 4.0 P@1), and no Tev-Gym environment shares a 13-gram or 8-gram with any evaluation task.

Bar chart of TB 4.0-CM Pass@1 for every model, highest first: Claude Opus 5 0.52, Claude Fable 5 0.46, GPT-5.6 (sol) 0.39, Tev-320B 0.37 (+28% over its base GLM-5.3-Flash, 0.29), Grok 4.6 0.31, Claude Opus 4.8 0.27, GPT-5.6 (terra) 0.23, Grok 4.5 0.20, Claude Sonnet 5 0.19, GPT-5.6 (luna) 0.18, Tev-27B 0.10 (2.5 times its base Qwen3.8-27B, 0.04), Gemini 3.7 Flash 0.09, Gemini 3.6 Flash 0.09, GLM 5.2 0.03.
All models on the Terminal-Bench 4.0 “CM” subset (TB 4.0-CM), ranked by Pass@1. Tev-320B and Tev-27B are trained with our recipe from GLM-5.3-Flash and Qwen3.8-27B, respectively; each arrow gives the relative gain over its base. *Evaluated by us, not on the official leaderboard.

Long horizon, in domain. Tev-27B raises P@5 from 0.17 to 0.28, matching GPT-6 Astra and Claude Opus 5, and closes more than two-thirds of the gap to Claude Fable 5 (0.33). It produces 33% more full solutions (P@1) than its base, and beats Qwen3.8-2.4T-A95B under the same harness on P@1. Tev-9B nearly doubles its base's best-of-5 solve rate.

ModelHarnessMath & ScienceMLCoding & SystemAll (Avg)
S@1S@5P@1P@5S@1S@5P@1P@5S@1S@5P@1P@5S@1S@5P@1P@5
Frontier API models
GPT-6 AstraCodex0.580.580.380.380.480.490.200.200.810.850.200.200.620.630.280.28
Claude Fable 5Claude Code0.470.510.380.380.520.570.240.400.830.900.200.200.580.630.290.33
GPT-5.6 (sol)Codex0.500.510.380.380.460.490.200.200.820.850.200.200.580.600.280.28
GPT-5.6 (terra)Codex0.510.560.380.380.250.260.000.000.750.830.200.200.500.550.220.22
Claude Opus 5Claude Code0.510.530.380.380.430.460.200.200.850.870.200.200.580.600.280.28
Claude Sonnet 5Claude Code0.500.510.380.380.390.560.120.400.800.870.200.200.550.620.260.33
Claude Opus 4.8Claude Code0.480.510.380.380.310.330.000.000.800.860.200.200.520.550.220.22
Gemini 3.7 FlashGemini CLI0.480.500.380.380.260.290.000.000.840.860.200.200.520.540.220.22
Gemini 3.6 FlashGemini CLI0.480.510.350.380.250.260.000.000.810.860.200.200.510.540.210.22
Gemini 3.1 ProGemini CLI0.490.520.350.380.270.290.000.000.820.880.200.200.520.560.210.22
Open-source models
Qwen3.8-2.4T-A95BVanillux20.200.380.200.380.100.200.040.200.740.860.160.200.320.460.140.28
DeepSeek-V4-proVanillux20.120.250.120.250.100.200.040.200.730.880.160.200.290.410.110.22
Qwen3.5-9B (base)Vanillux20.070.120.000.000.190.260.000.000.590.730.200.200.250.330.060.06
Tev-9B (ours)Vanillux20.140.290.040.120.230.260.000.000.560.730.200.200.290.400.070.11
Δ (Tev-9B − Qwen3.5-9B)–0.07↑0.17↑0.04↑0.12↑0.04↑–––0.03↓–––0.04↑0.07↑0.01↑0.05↑
Qwen3.8-27B (base)Vanillux20.150.250.150.250.210.320.000.000.740.820.200.200.330.430.120.17
Tev-27B (ours)Vanillux20.230.500.230.500.260.290.000.000.820.860.200.200.400.540.160.28
Δ (Tev-27B − Qwen3.8-27B)–0.08↑0.25↑0.08↑0.25↑0.05↑0.03↓––0.08↑0.04↑––0.07↑0.11↑0.04↑0.11↑

S@k is the mean graded score and P@k the fraction of tasks fully solved: @1 averages 5 rollouts per task, @5 is best-of-5.

Analysis

What matters: the latent world and the dense reward

MethodDataRewardGPU-h
/ step
ALE-CM
WorldDAGHidden truthSignalCreditAdv. norm.S@1S@5P@1P@5
Qwen3.5-9B (base)–––no RL–––0.250.330.060.06
Tev-9B (ours)✓✓✓ΔΦt + terminalGt, γ=1per turn9.00.290.400.070.11
Data ablation: replace one construction step with an LLM; terminal reward, compare with Sparse
Public-view verifier✓✓LLMterminalwhole rolloutper rollout22.30.240.350.040.06
Prompt-firstLLMLLMLLMterminalwhole rolloutper rollout11.80.230.320.060.06
Reward ablation: keep Tev-Gym, change how Φ enters training
Sparse (terminal only)✓✓✓terminalwhole rolloutper rollout16.70.250.320.060.06
PPRS, pooled norm.✓✓✓ΔΦt + terminalGt, γ=1per rollout10.90.280.350.060.06
Direct reward (additive)✓✓✓Φ(st) + terminalGt, γ=0per turn9.70.290.350.060.06

Ablations on ALE-CM with Qwen3.5-9B. In the data ablations, GPT-5.6 Sol (xhigh) replaces the marked construction steps. GPU-h/step is the mean time of one training step times the recipe's 32 GPUs, in H100 GPU-hours.

The latent world is what makes the data work

Letting an LLM write the verifier from the agent's view (public-view verifier) or generate the whole task prompt-first leaves no hidden truth to check against. Neither improves over the base model in S@1 (0.24 and 0.23 vs. 0.25), and the public-view variant discards 88% of sampled groups during training. The full method leads both by 0.05–0.06 S@1 and 0.05 P@5.

Dense reward trains better and cheaper

On the same data, PPRS improves S@1 by 0.04 over terminal-only reward at 9.0 instead of 16.7 GPU-hours per step. Per-turn normalization keeps credit for later turns, and \(\gamma = 1\) carries progress back to earlier actions. At 27B the gap widens: terminal-only reward yields almost entirely zero-advantage groups and never completes a single optimizer step.

Qualitative

Example: root-causing a crashing Kubernetes service

ALE-CM task k8s_payment_api_root_cause_analysis (Coding & System). From a kubectl capture, the Deployment/HPA YAML and a crashing pod's log, the agent writes a JSON root-cause report with verbatim evidence strings and a rollback verdict. Every agent finds the same root causes, but Tev-27B and Claude Fable 5 ground the liveness-probe cause in the measurement that explains it and judge rollback safe. The base model and GPT-6 Astra cite only the symptoms and decline rollback, forfeiting the decisive grader component.

Claude Fable 50.97
t3Plans: read, write, validate …
t4–10Reads prompt, cluster state, YAML, pod log …
t14Writes the report … probe cause quotes initialDelaySeconds: 5 … cache warm-up complete (duration=6.67s) … “takes roughly 8 s to start … probe fires before the server is listening” … safe_to_rollback: true … “pod-template settings of revision 17, not … schema changes.”
t18–21Verifies every quote is verbatim … summarises the incident … submits.
GPT-6 Astra0.75
t3Plans: “inspect the evidence, write the RCA, and validate its quotes” …
t6Dumps cluster state and YAML in one shell call … checks for apply_patch …
t8“OOM kills are confirmed. HPA values conflict; rollback safety is unverified” → rollback verdict: not safe …
t9Writes the report … probe cause cites the settings and failures but not the warm-up measurement … not credited.
t11–13Validates schema, 41 quotes … submits: “… why rollback safety remains unverified.”
Qwen3.8-27B base0.70
t2Lists the task directory …
t4–10Reads the four inputs …
t12–16Checks the em-dash and column spacing byte by byte … writes build_rca.py to extract exact lines …
t18Builds the report … probe cause: “the app needs ~6.6 s … for cache warm-up …” but evidence = probe-failure lines + YAML block only … “Do NOT roll back … it would OOM/crash-loop in the same way” → safe_to_rollback: false.
t20–24Independent re-verification: validity, schema, 48 quotes verbatim … submits.
Tev-27B (ours)0.97
t2–18Reads the four inputs … confirms planned quotes byte for byte …
t20Probe cause: initialDelaySeconds: 5 … cache warm-up complete (duration=6.67s …) … “warm-up … longer than the 5 s initial delay” …
t28Verifies 29 quotes …
t30Reconsiders rollback: “… ‘started flapping after a routine chart-value change’ … roll back first” → safe_to_rollback: true.
t32–40Re-validates … submits.
decision that earns a grader componentdecision that forfeits one
Cite

BibTeX

@misc{hu2026tev,
  title  = {{Tev}: Scaling Long-Horizon Agentic Environments with Dense Progress Rewards},
  author = {Hu, Wenbo and Pramanick, Shraman and Singh, Krishna Kumar and
            Peng, Nanyun and Chang, Kai-Wei and Lee, Yong Jae},
  year   = {2026},
  note   = {Preprint}
}