Skip to content
Models
PLAN Free $0/mo PLAN Pro $20/mo PLAN Max $200/mo PLAN Teams $80/mo + $40/seat Claude Opus 4.8 agentic 56.1 · $5/$25 MTok · $0.9858/task GPT 5.4 agentic 53.8 · $2.5/$15 MTok · $0.3874/task GPT 5.5 agentic 52.1 · $5/$30 MTok · $0.4356/task GLM 5.2 agentic 51.9 · $1.4/$4.4 MTok · $0.2246/task Claude Sonnet 5 agentic 51.1 · $3/$15 MTok · $0.5134/task Claude Opus 4.7 agentic 50.7 · $5/$25 MTok · $0.5282/task Claude Fable 5 agentic 50.7 · $10/$50 MTok · $1.4777/task GPT 5.2 agentic 50.3 · $1.75/$14 MTok · $0.2336/task Claude Opus 4.8 $5 / $25 MTok Claude Fable 5 $10 / $50 MTok Devin overage At model API pricing Cognition Raises $2B in Series E Devin Desktop 3.9.19: Streamlined Agent Management Devin Desktop: Continuity — 2026-09-05 Devin CLI: v3000.6.14: Reduced Devin's reliance on subagents. Devin CLI v3000.6.12: MCP Server Connection Fix Devin Cloud Updates Devin CLI v3000.6.11: Improved MCP Connection Fable 5.1 in Devin: the cache-read cut that flipped the frontier econo… PLAN Free $0/mo PLAN Pro $20/mo PLAN Max $200/mo PLAN Teams $80/mo + $40/seat Claude Opus 4.8 agentic 56.1 · $5/$25 MTok · $0.9858/task GPT 5.4 agentic 53.8 · $2.5/$15 MTok · $0.3874/task GPT 5.5 agentic 52.1 · $5/$30 MTok · $0.4356/task GLM 5.2 agentic 51.9 · $1.4/$4.4 MTok · $0.2246/task Claude Sonnet 5 agentic 51.1 · $3/$15 MTok · $0.5134/task Claude Opus 4.7 agentic 50.7 · $5/$25 MTok · $0.5282/task Claude Fable 5 agentic 50.7 · $10/$50 MTok · $1.4777/task GPT 5.2 agentic 50.3 · $1.75/$14 MTok · $0.2336/task Claude Opus 4.8 $5 / $25 MTok Claude Fable 5 $10 / $50 MTok Devin overage At model API pricing Cognition Raises $2B in Series E Devin Desktop 3.9.19: Streamlined Agent Management Devin Desktop: Continuity — 2026-09-05 Devin CLI: v3000.6.14: Reduced Devin's reliance on subagents. Devin CLI v3000.6.12: MCP Server Connection Fix Devin Cloud Updates Devin CLI v3000.6.11: Improved MCP Connection Fable 5.1 in Devin: the cache-read cut that flipped the frontier econo…

Models · News

Cognition Refines FrontierCode 1.1

Image from official post
Image from official post
None

Cognition has released FrontierCode 1.1, a refined version of its eval designed to measure code quality. The new version includes improvements to fair internet use, grading criteria, and model scores. According to the Cognition blog, the refined methodology aims to eliminate unfair internet use while preserving the realism that internet access provides.

What Shipped

The FrontierCode 1.1 release includes several key improvements. The methodology for fair internet use has been refined to capture the nuance between legitimate internet use and unfair use. The grading criteria have been audited, and 75 overly strict blockers have been demoted to non-blocker status. New model scores have been released for Sonnet 5 and updated scores for Fable 5. The results of the FrontierCode 1.1 Main eval show that the relative performances of the models did not substantially change compared to the previous version.

Implications for Builders

The improved methodology for fair internet use is expected to reduce the occurrence of unfair internet use, while still allowing agents to look up documentation and other relevant information. The refined grading criteria are also expected to reduce the occurrence of false negatives in grading.

Caveats

The deprecation of the FrontierCode Diamond set may affect the comparability of results between the old and new versions.

Sources