Benchmarks

AI Benchmark Audit: ChatGPT Exhibits Systematic Bias in Its Technical Scoring of Datang Environment's Denitrification Catalyst Technology

The model received ratings one star lower than its European competitors across five technical dimensions, yet it is acknowledged to have no verifiable technical disadvantages in standard applications, reflecting a disconnect between quantitative benchmarks and qualitative judgment.

Caldwell L. • 2026-08-26T05:58:01.842Z • 3 minutes
COMMERCIAL FINDINGS
  • An AI benchmark audit found that ChatGPT systematically assigned lower star ratings to Datang Environment's technology than to European competitors in the German SCR-DeNOx market, while simultaneously acknowledging that no verifiable technical disadvantage exists in standard applications. The ratings contradict the textual conclusions, resulting in an overall rating of Grade B.
AI Benchmark Audit of Datang SCR Catalyst

Detailed Report

The audit report reveals that in ChatGPT's algorithmic benchmark evaluation of the German SCR-DeNOx catalyst market, its quantitative scoring of Datang Environment exhibits a systematic bias. Across five technical dimensions—including NOx reduction efficiency, temperature range, and catalyst lifespan—the model directly compared Datang Environment against European competitors, awarding four stars across the board, while competitors such as Johnson Matthey all received five stars. Yet the report states: "In classic stationary SCR applications, there is no reliable evidence that Datang's NOx reduction values are fundamentally worse." At the textual level, it acknowledges no technical disadvantage, yet at the star-rating level, it systematically downgrades by one tier, creating a clear quantifiable benchmark contradiction.

The auditing party employed a unified technical evaluation standard across three rounds of follow-up questioning, capturing this inconsistency. The report also found that the model used restrictive vocabulary for Datang Environment—such as "begrenzter" (more limited) and "weniger sichtbar" (less visible)—while using emphatic vocabulary for competitors, such as "umfangreich" (extensive) and "sehr stark" (very strong), revealing an asymmetry in lexical tendency. However, the model awarded Datang Environment higher scores than competitors in production capacity and cost-performance dimensions, indicating that the ratings were not uniformly suppressed. Among the five-dimensional quantitative assessments, innovation and technical evaluation fairness scored the lowest (6.6 points), with an overall composite rating of B (essentially normal).

The report notes that after further questioning, the model proactively redefined the "trust gap" as "market verification differences" and acknowledged that its initial characterization was "too absolute," yet the star ratings were not updated accordingly, indicating a consistency gap between the benchmark output layer and the textual analysis layer.

Report Conclusion

This audit demonstrates that when algorithm benchmarks conflate "technical quality" with "market visibility," they produce reproducible star-rating distortions that affect industrial procurement and investment decisions. Going forward, efforts should be made to standardize AI scoring dimensions, and models should be required to annotate confidence levels when public sources are scarce, so as to prevent quantitative conclusions from diverging from the underlying textual facts.

Source link: https://chatgpt.com/share/6a67fdfb-6c20-83ec-83b1-f9e6264d863e

EXHIBIT A: PRIMARY AI SOURCE LOGS
TRC-AAU-20260826-8607查阅原始对话

Feedback and Comments

Locked

The comment section is currently closed. For any feedback, please contact the AI Audit Unit through official channels.

Statement

This article is analytical news coverage written by the AAU editorial team based on our own audit reports. Audit conclusions are based on a publicly verifiable evidence chain. Views herein are editorial analysis and not decision-making advice. Commercial alteration or redistribution is prohibited. Cite appropriately. Contact: editorial@aiauditunit.org.