Skip to content

Frontier AI Risk Monitor update: risk indices rise severalfold as multiple models cross key threshold

On July 19, 2026, Concordia AI launched the Frontier AI Risk Monitoring Platform v2.0 at the World Artificial Intelligence Conference (WAIC). The release includes three major updates: the Risk Index v2.0, the 2026 Q2 Risk Monitoring Report, and the Frontier AI Risk Benchmark Database. This article focuses on the findings of the 2026 Q2 report while briefly introducing the upgraded Risk Index and the new Benchmark Database supporting it.

A More Realistic View of AI Risk

This is the first report to use the Risk Index v2.0 framework. The upgrade is designed to answer two simple questions more clearly: how dangerous could a model become as its capabilities grow, and how much of that danger remains after its safeguards are taken into account?

Compared with the previous version, v2.0 better reflects the possibility that risk can accelerate as models become more capable. It also looks beyond whether a model refuses an ordinary harmful request by testing how well safeguards hold up against jailbreaks or deliberate tampering. Two “Yellow Lines” make the results easier to interpret: the Capability Yellow Line signals that a model could significantly increase severe-harm risks without safeguards, while the Risk Yellow Line signals that high risk remains even with existing safeguards in place.

The underlying evaluations have also been refreshed with more realistic tasks and advanced jailbreak testing, helping the Platform track both what frontier models can do and where their protections may fail.

The report covers 47 frontier models released by 13 leading AI companies from 2025 Q3 to 2026 Q2, including GPT-5.5, Claude Opus 4.8, Grok 4.3, DeepSeek V4 Pro, Qwen 3.7 Max, Doubao Seed 2.1 Pro, HY 3.0 Preview, ERNIE 5.1, MiniMax M3, Kimi K2.6, GLM 5.2, and MiMo V2.5 Pro.

Key Findings

The report highlights seven key findings:

 

1. Cyber, biological, and loss-of-control Risk Indices have risen severalfold in less than a year, with multiple models in each domain now exceeding the Capability Yellow Line

Under Risk Index v2.0, the average Risk Index of evaluated models rose rapidly in less than a year: 4.4x in cyber offense, 6.6x in biological risks, and 2.4x in loss-of-control.

Multiple models have already crossed the Capability Yellow Line: 4 models in cyber offense, 22 in biological risks, and 12 in loss-of-control. Crossing this line means that, without safeguards, a model would significantly increase severe-harm risk relative to non-AI baselines.

Note: 1. In the current v2.0 framework, Risk Indices have not yet been calculated for chemical risks or harmful manipulation, so this section focuses on cyber offense, biological risks, and loss-of-control.
2. Because Risk Index v2.0 makes substantial changes to both the benchmark suite and metric calculation, values in this report should not be compared item by item with those from v1.0 or v1.5. They are better used to compare models, quarters, and domains within the v2.0 framework.

2. Model risk profiles are diverging: Gemini 3.1 Pro Preview has the highest biological and loss-of-control Risk Indices, while DeepSeek V4 Pro shows more pronounced cyber misuse risk

Gemini 3.1 Pro Preview has the highest overall Risk Indices in biological risks and loss-of-control, and both have crossed the Risk Yellow Line. Unlike the Capability Yellow Line, crossing the Risk Yellow Line means that, even with existing safeguards, the model’s residual risk still reaches the high-risk boundary.

DeepSeek V4 Pro has the highest Risk Index in cyber offense, and its loss-of-control Risk Index is second only to Gemini 3.1 Pro Preview.

The GPT family is rising quickly in cyber offense and loss-of-control, and has crossed the Risk Yellow Line in biological risks.

Kimi, MiniMax, Qwen, Doubao, and other families are also rising quickly in cyber offense and biological risks; in biological risks, they have already crossed the Risk Yellow Line.

Note: This evaluation does not include Claude’s strongest models, Fable/Mythos 5.

3. Proprietary models continue to dominate the capability frontier across multiple domains, while open-weight models are generally weaker on safeguards for cyber, biological, chemical, and manipulation risks

Proprietary models still dominate the capability frontier across domains, with GPT-5.5, Claude Opus 4.8, Gemini 3.1 Pro Preview, and other proprietary models maintaining a capability lead in multiple domains.

Across the four misuse-risk domains of cyber offense, biological risks, chemical risks, and harmful manipulation, open-weight models have substantially lower Safety Scores than proprietary models.

In loss-of-control, Safety Score distributions are similar for open-weight and proprietary models, but some proprietary models, such as Gemini 3.1 Pro Preview and Grok 4, have notably lower Safety Scores. These models pose higher risk because they combine high Capability Scores with low Safety Scores.

4. Leading models’ long-horizon vulnerability exploitation and penetration capabilities are improving quickly, while safeguards such as prompt-injection defenses are regressing

In less than a year, the top model score rose by 68% on CyBench and 35% on CVE-Bench, indicating rapid progress by frontier models on autonomous, long-horizon vulnerability exploitation and penetration tasks.

At the same time, cyber safeguards have not improved in step. The average Safety Score fell by 3% in the latest quarter, and latest-quarter models’ average performance on prompt-injection defense (PromptInjection) regressed.

5. More than ten models now exceed human-expert levels in wet-lab troubleshooting and sequence understanding, but basic safeguards such as biological refusal have not improved in step

On capabilities, 22 models now exceed human-expert levels in wet-lab troubleshooting (BioLP-Bench), and 11 models exceed human-expert levels in sequence understanding (SeqQA). Frontier models’ biological reasoning (FrontierScience-Bio) and bioinformatics capabilities (BixBench) continue to improve.

On safety, the average Safety Score is broadly unchanged from the previous quarter, and basic biological refusal (BiologicalHarmfulQA) showed no clear improvement in the latest quarter.

6. The latest-quarter models did not set a new high in loss-of-control risk, but weak honesty and covert influence over users still reveal safety gaps

Compared with the rapid growth seen last quarter, models released in the latest quarter did not set a new high in the loss-of-control Risk Index, and none reached the Risk Yellow Line.

On capabilities, self-replication (Self-Proliferation), self-improvement (MLE-Bench), situational awareness (SAD-mini), and related capabilities did not set new highs.

On safety, loss-of-control safeguards still show no clear improvement: model honesty (MASK) remains weak overall, and some models still show a pronounced tendency to covertly influence users (DarkBench).

7. Under advanced jailbreak attacks, average safety falls sharply across all misuse-risk domains

After jailbreak red-teaming attacks are added, frontier models’ average safety scores fall from 78.2 to 8.9 in biological risks, from 90.5 to 29.2 in cyber offense, from 83.7 to 53.1 in chemical risks, and from 87.8 to 36.4 in harmful manipulation.

Jailbreak resistance varies greatly across models. Claude Opus 4.8 maintains an average refusal rate of 69.1% under red-team attacks, compared with just 5.8% for Hunyuan T1 (250711) under the same conditions.

Note: Jailbreak red-team attacks apply only to misuse-risk domains, not to loss-of-control risk.

See the full report for detailed monitoring results.

Recommendations for Stakeholders

Based on these findings, the report makes the following recommendations:

  • For model developers: Pay close attention to your models’ Risk Indices, especially whether they cross the Risk Yellow Line. For higher-risk models, prioritize improving Base Safety, Jailbreak Safety, and loss-of-control safety. If better safeguards alone cannot sufficiently mitigate risk, consider reducing high-risk capabilities and establish clear risk thresholds, mitigation measures, and release policies.
  • For AI safety researchers: Continue exploring more effective methods for capability elicitation, red-team attacks, prompt injection, and multi-turn manipulation to accurately assess the upper bounds of model capability and lower bounds of safety. At the same time, develop more effective model hardening, dangerous-capability removal, and risk mitigation approaches suited to open-weight models.
  • For policymakers: The report identifies early warning signs in cyber offense, biological risks, and loss-of-control. Strengthen ongoing monitoring and risk analysis of pre-release evaluations and mitigations, and adopt differentiated governance based on model capability, safety, and open-weight or proprietary distribution.

Concordia AI will continue using the Frontier AI Risk Monitoring Platform to track risk trends among the world’s most advanced AI models and provide data to support the development of safe and trustworthy AI.

A Searchable Database for AI Risk Evaluations

Concordia AI also launched the Frontier AI Risk Benchmark Database, bringing together more than 200 benchmarks—standardized evaluations used to test AI systems—published between 2023 and 2026. Information that was previously scattered across research papers, code repositories, datasets, and model reports can now be searched and compared in one place.

The Database organizes benchmarks by risk area, evaluation purpose, and availability. It also connects them to real-world risk scenarios, helping users understand not only what a benchmark measures, but how it contributes to a broader assessment of AI risk. This provides a more transparent foundation for the monitoring results in this report and a practical resource for researchers, model developers, and policymakers.

Explore the Benchmark List and Risk Model.

References

Back To Top