Open Access Article
Social Development Review DOI: DOI:10.64635/ja.2026.1447.
*通讯作者: 无
发布时间: 2026-08-22 总浏览量: 25
训练数据的样本失衡、历史偏见、标签噪声和代理变量关联会被人工智能模型学习并放大,仅依赖准确 率或单一公平指标难以及时发现偏差。针对该问题,提出一种面向模型训练前审计的多维训练数据偏差检测方法。 该方法从群体代表性、特征分布、标签一致性、群体公平和反事实敏感性五个维度构建指标,经过方向统一、标准 化与权重学习形成综合偏差得分,同时输出偏差群体、关联特征和疑似异常标签。采用包含1200个数据集窗口的 可控仿真进行验证,每个窗口含600个样本,并注入代表性偏差、特征偏移、标签偏差、代理偏差和混合偏差。结 果表明,所提方法的精确率、召回率、F1值和AUC分别达到0.953、0.971、0.962和0.993,对多类型偏差的识别 能力优于样本比例、统计均等和敏感系数等单指标方法。
Sample imbalance, historical bias, label noise, and proxy-variable associations in training data can be learned and amplified by artificial intelligence models, and reliance on accuracy or a single fairness metric alone makes it difficult to detect bias in a timely manner. To address this problem, this paper proposes a multidimensional training data bias detection method for pre-training model audits. The method constructs indicators from five dimensions: group representativeness, feature distribution, label consistency, group fairness, and counterfactual sensitivity. Through direction unification, standardization, and weight learning, a comprehensive bias score is formed, while biased groups, associated features, and suspected abnormal labels are also output. Validation is conducted using a controllable simulation containing 1,200 dataset windows, each with 600 samples, with representativeness bias, feature shift, label bias, proxy bias, and mixed bias injected. The results show that the proposed method achieves a precision, recall, F1 score, and AUC of 0.953, 0.971, 0.962, and 0.993, respectively, and outperforms single-indicator methods such as sample proportion, statistical parity, and sensitivity coefficient in identifying multiple types of bias.
[1]陈晋音,陈奕芃,陈一鸣,等.面向深度学习的公平性研究综述[J].计算机研究与发展,2021,58(2):264-280.
[2]古天龙,李龙,常亮,等.公平机器学习:概念、分析与设计[J].计算机学报,2022,45(5):1018-1051.
[3]范卓娅,孟小峰.算法公平与公平计算[J].计算机研究与发展,2023,60(9):2048-2066.
[4]王艳,侯哲,黄滟鸿,等.基于概率模型检查的树模型公平性验证方法[J].软件学报,2022,33(7):2482-2498.
[5]陈彩华,佘程熙,王庆阳.可信机器学习综述[J].工业工程,2024,27(2):14-26.
[6] 彭丽徽,张琼,李天一.人工智能嵌入政府数据治理的算法歧视风险及其防控策略研究[J].农业图书情报学报,2024.