
Dissolved gas analysis (DGA) is widely used for condition assessment of oil-immersed power transformers, but public-data artificial intelligence (AI) studies can overstate evidential reliability when labels are heterogeneous, samples overlap across benchmarks, or random splits ignore asset and time structure. This study presents an evidence-hierarchy audit framework for trustworthy AI-assisted transformer DGA assessment. The framework separates public-label benchmark agreement, transparent weak-label reproduction, leakage-sensitive validation, and future field-confirmed validity; combines exact gas-label fingerprinting, tree-ensemble baselines, Shapley additive explanations (SHAP) attribution, and a Mechanism-Consistency Score (MCS); and uses confidence-MCS logic as an expert-review trigger. On the Kaggle five-gas benchmark, LightGBM, Random Forest, and XGBoost achieved mean Macro-F1 values of 0.8982, 0.8884, and 0.8878, while exact fingerprinting showed that all 2321 alan-456 samples overlapped with Kaggle-derived fingerprints. After overlap removal, cross-dataset Macro-F1 decreased to 0.4688–0.5015. On the Lewis weak-label data, transformer-grouped validation reduced Macro-F1 from 0.9992 to 0.8691, showing sensitivity to asset-level splitting assumptions. MCS analysis identified class-dependent mechanism consistency and low-MCS review signals under five-gas inputs. The results support public-benchmark reliability and evidence-quality auditing rather than field-confirmed diagnostic performance.
dissolved gas analysis; power transformer; trustworthy artificial intelligence; weak labels; benchmark reliability; SHAP; Mechanism-Consistency Score