The headline is simple: model release order did not predict engineering precision. The same sealed questions and scoring rules produced a stable leader—and a revealing gap between selecting an answer and supporting it correctly.
The updated leaderboard
| Rank | Model | Overall | Answers | Standards |
|---|---|---|---|---|
| 1 | DeepSeek V4 Pro | 97.8% | 99.4% | 91.7% |
| 2 | GPT-5.6 Sol | 97.5% | 99.7% | 88.9% |
| 3 | Kimi K3 | 97.1% | 99.4% | 88.0% |
| 4 | Grok 4.5 | 96.7% | 99.4% | 86.1% |
| 5 | Claude Sonnet 5 | 96.2% | 98.8% | 86.1% |
| 6 | Grok 4.6 | 95.7% | 98.1% | 86.1% |
| 7 | Gemini 3.5 Flash | 95.1% | 97.8% | 84.3% |
| 8 | Qwen3.8 Max | 94.6% | 98.1% | 80.6% |
Grok 4.6 did not beat Grok 4.5
Grok 4.6 finished at 95.7%, one point behind Grok 4.5 at 96.7%. Their standards citation scores were identical at 86.1%; the difference came from answer accuracy, where Grok 4.6 scored 98.1% against Grok 4.5's 99.4%.
That is exactly why independent, task-specific evaluation matters. A newer model may be stronger in broad capability tests and still regress on a narrow professional protocol.
Kimi K3 was the strongest of the new generation
Kimi K3 placed third overall at 97.1%. It matched the field's best answer accuracy tier at 99.4% and delivered an 88.0% standards score. Its 97.6% standards-task score made it the most convincing new entrant when the question required both an engineering answer and the governing code or standard.
Qwen3.8 Max exposed the reliability tax
Qwen3.8 Max selected answers well—98.1% were correct—but its standards citation F1 fell to 80.6%. 4 of its 320 responses were still unusable after the benchmark's retry policy. We retained those failures as zeroes, because silently removing them would overstate deployment reliability.
The real separator was standards discipline
The new models all remained excellent on engineering analysis. The larger spread appeared when they had to name the correct governing standard without adding unsupported citations. This is a useful warning for real practice: polished reasoning and a correct multiple-choice selection do not guarantee dependable code grounding.
What this benchmark does not prove
PE Precision Bench is a one-run, closed-book, zero-temperature multiple-choice protocol. It does not establish engineering competence and it does not replace current standards, project documents, or review by a licensed engineer. It measures a narrower question: under identical constraints, which models most precisely combine answer selection with standards support?
