Bro
Bro Evals
13 pages · / search · ? help

Bro Frontier Index 1.2

Practical-case scores across nine categories. Higher is better.

Models
preliminary set
Frontier
Best open

Model performance

Capability, coverage, and operating economics in one grid.

50+30–49<30
Preliminary scores renormalize over measured categories.Estimated entries are tagged. The leaderboard has its own estimates toggle.Cyber in v1.2 is CY1, CY3 and CY4. K4 and AIML2 start as TBD.

Task weights

How categories combine into the overall score.

Legacy benchmark

Bro Capabilities Index

The original seven-category capability benchmark, kept for historical comparison.

Models
parsed entries
Frontier
Best open

BCI archive

Overall plus seven category averages.

Security benchmark

Bro Cyber Eval

Security-focused evaluation over CY1 to CY4.

Models
measured
Frontier
Best open

Cyber performance

Task-level scores and the weighted overall result.

Model against model

Compare models

Line up four models across every metric.

Benchmark
BFI
active view
Models
side by side
Source
evals-clean.txt
same parsed data
Benchmark