Skip to content
Models
PLAN Free $0/mo PLAN Pro $20/mo PLAN Max $200/mo PLAN Teams $80/mo + $40/seat Claude Opus 4.8 agentic 56.1 · $5/$25 MTok · $0.9858/task GPT 5.4 agentic 53.8 · $2.5/$15 MTok · $0.3874/task GPT 5.5 agentic 52.1 · $5/$30 MTok · $0.4356/task GLM 5.2 agentic 51.9 · $1.4/$4.4 MTok · $0.2246/task Claude Sonnet 5 agentic 51.1 · $3/$15 MTok · $0.5134/task Claude Opus 4.7 agentic 50.7 · $5/$25 MTok · $0.5282/task Claude Fable 5 agentic 50.7 · $10/$50 MTok · $1.4777/task GPT 5.2 agentic 50.3 · $1.75/$14 MTok · $0.2336/task Claude Opus 4.8 $5 / $25 MTok Claude Fable 5 $10 / $50 MTok Devin overage At model API pricing Cognition Raises $2B in Series E Devin Desktop 3.9.19: Streamlined Agent Management Devin Desktop: Continuity — 2026-09-05 Devin CLI: v3000.6.14: Reduced Devin's reliance on subagents. Devin CLI v3000.6.12: MCP Server Connection Fix Devin Cloud Updates Devin CLI v3000.6.11: Improved MCP Connection Fable 5.1 in Devin: the cache-read cut that flipped the frontier econo… PLAN Free $0/mo PLAN Pro $20/mo PLAN Max $200/mo PLAN Teams $80/mo + $40/seat Claude Opus 4.8 agentic 56.1 · $5/$25 MTok · $0.9858/task GPT 5.4 agentic 53.8 · $2.5/$15 MTok · $0.3874/task GPT 5.5 agentic 52.1 · $5/$30 MTok · $0.4356/task GLM 5.2 agentic 51.9 · $1.4/$4.4 MTok · $0.2246/task Claude Sonnet 5 agentic 51.1 · $3/$15 MTok · $0.5134/task Claude Opus 4.7 agentic 50.7 · $5/$25 MTok · $0.5282/task Claude Fable 5 agentic 50.7 · $10/$50 MTok · $1.4777/task GPT 5.2 agentic 50.3 · $1.75/$14 MTok · $0.2336/task Claude Opus 4.8 $5 / $25 MTok Claude Fable 5 $10 / $50 MTok Devin overage At model API pricing Cognition Raises $2B in Series E Devin Desktop 3.9.19: Streamlined Agent Management Devin Desktop: Continuity — 2026-09-05 Devin CLI: v3000.6.14: Reduced Devin's reliance on subagents. Devin CLI v3000.6.12: MCP Server Connection Fix Devin Cloud Updates Devin CLI v3000.6.11: Improved MCP Connection Fable 5.1 in Devin: the cache-read cut that flipped the frontier econo…

Models · News

Cognition Measures Trust in Open-Source Models

Image from official post
Image from official post
With targeted post-training and other techniques, open-source models can be made at least as safe as leading closed models.

Cognition, a company developing autonomous software engineers, has created an evaluation suite to assess the trustworthiness of open-source-derived models. The suite combines direct questioning to check for propaganda output and realistic coding scenarios to ensure model behavior remains constant across users and contexts.

What Shipped

The evaluation suite was run on a range of models, including Kimi K2.7 Code, the open-source base model from which SWE-1.7 was developed. The results indicate that SWE-1.7 performs as well or better on the trustworthiness evaluation suite than models from leading U.S.-based frontier labs. The trust evaluations come in two parts: propaganda and censorship, and security and vulnerabilities. The evaluations probed models with 145 politically sensitive questions and graded every response along six axes, such as active propaganda rate and factual accuracy.

Implications for Builders

The results suggest that open-source models are not inherently unsafe, and with targeted post-training and other techniques, they can be made at least as safe as leading closed models. This has implications for builders who rely on open-source models, as they can now develop new models from open-source starting points with increased confidence in their trustworthiness.

Caveats

However, the company notes that they are still actively developing and building their trustworthiness evaluation suite, and the initial results are subject to further refinement.

Sources