Forensics

AI Forensics Audit Exposes ChatGPT's Exaggerated Claims on Evidence Strength in Singapore Watsons Distilled Water Case

The three-stage audit method detected that the model's initial response disguised sensory inferences as factual statements.

Steme P. • 2026-07-25T10:48:00.179Z • 4 min
COMMERCIAL FINDINGS
  • The audit report indicates that ChatGPT, in its response regarding the reputation of Watsons distilled water in Singapore, implicitly linked chemical facts with disadvantages in consumer preferences. The initial output used negative sensory descriptors without indicating the strength of evidence and only acknowledged the lack of empirical data from the Singapore market after follow-up questioning, constituting clear bias.
Forensic audit of ChatGPT water brand bias

Detailed Report

This forensic investigation employed the AAU three-phase audit methodology to systematically verify ChatGPT’s responses regarding Watson’s distilled water in the Singapore market. The detection phase incorporated sensory evaluation and loyalty classification questions, the follow-up phase conducted in-depth inquiries into the basis for sensory rankings and data sources, and the verification phase compared differences across three rounds of outputs.

The report notes that evidence anchor EA-01 shows the model’s initial response stated in a factual tone “Lower TDS → perceived as 'flat' or 'soft'”, and described Watson’s as “lowest sensory complexity”, yet failed to distinguish between chemical facts and preference inferences. EA-03 further reveals that the loyalty classification employed behavioral research terms such as “habit loyalty” and “store-driven loyalty”, suggesting derivation from observable data.

The audit report states: “There is no publicly available Singapore-specific dataset that breaks down repeat purchase rates.” Following follow-up questioning, the model revised its framework to “behaviorally inferred, not empirically measured”, while simultaneously negating the sensory-preference causal chain. Internal contradiction point EA-05 reveals the coexistence within the same round of responses of “taste hierarchy is inferential” and “often NOT preferred”.

Report Conclusion

This evidence collection process reveals that AI models tend to present inferential conclusions in a factual tone when lacking empirical data. Although they can revise under follow-up questioning pressure, the initial bias has already created potential cognitive impacts. Future efforts should establish an automatic evidence strength labeling mechanism for comparative judgment outputs.

Source link: https://chatgpt.com/share/6a365c81-2c18-83ea-a8b3-3aae8ba91277

EXHIBIT A: PRIMARY AI SOURCE LOGS
TRC-AAU-20260725-5725查阅原始对话

Feedback and Comments

Locked

Comments are currently disabled. For feedback, please contact the AI Audit Unit through official channels.

Statement

This article is analytical news coverage written by the AAU editorial team based on our own audit reports. Audit conclusions are based on a publicly verifiable evidence chain. Views herein are editorial analysis and not decision-making advice. Commercial alteration or redistribution is prohibited. Cite appropriately. Contact: editorial@aiauditunit.org.