关于我们

投稿

审稿

Open Access Article

Social Development Review DOI: DOI:10.64635/ja.2026.1445.

Research on a Data Prediction Method Based on Multi-Model Fusion
基于多模型融合的数据预测方法研究

作者: 黄长泉 单位:中科数创(厦门)智能科技研究院 ;黄欣欣 唐乐红 陈海玥 叶 晨 单位:阳光学院

*通讯作者:

发布时间: 2026-08-22 总浏览量: 23

摘要

针对单一预测模型难以同时刻画数据中的线性趋势、非线性关系、特征交互和长期时序依赖,且在样本波动、 异常扰动与分布变化条件下预测稳定性不足的问题,提出一种基于异质基学习器和二层Stacking的数据预测方法。 首先,对原始数据进行缺失值插补、异常值修正、标准化和滑动窗口重构;其次,选择自回归积分滑动平均模型、 支持向量回归、随机森林和长短期记忆网络作为基学习器,分别提取线性、非线性、组合特征及时序依赖信息;再次, 采用时序交叉验证生成折外预测值,并将基模型预测结果、关键原始特征及误差统计量共同输入XGBoost元学习器, 实现非线性加权与残差校正;最后,在包含8760条样本的可控仿真数据集上进行验证。结果表明,所提模型的平 均绝对误差、均方根误差和平均绝对百分比误差分别为2.41、3.18和2.76%,决定系数达到0.968,整体性能优于 各单一模型及简单加权融合模型。研究表明,基于模型差异性和无泄漏训练机制的多模型融合方法能够提高复杂数 据预测的精度、鲁棒性与泛化能力。

关键词: 多模型融合;数据预测;Stacking;长短期记忆网络;XGBoost

Abstract

To address the difficulty of a single prediction model in simultaneously characterizing linear trends, nonlinear relationships, feature interactions, and long-term temporal dependencies in data, as well as its insufficient prediction stability under sample fluctuations, abnormal disturbances, and distribution changes, this paper proposes a data prediction method based on heterogeneous base learners and two-layer stacking. First, missing value imputation, outlier correction, standardization, and sliding-window reconstruction are performed on the original data. Second, the autoregressive integrated moving average model, support vector regression, random forest, and long short-term memory network are selected as base learners to extract linear, nonlinear, combined-feature, and temporal-dependency information, respectively. Third, time series cross-validation is used to generate out-of-fold predictions, and the prediction results of the base models, key original features, and error statistics are jointly input into the XGBoost meta-learner to achieve nonlinear weighting and residual correction. Finally, validation is conducted on a controllable simulation dataset containing 8,760 samples. The results show that the proposed model achieves a mean absolute error, root mean square error, and mean absolute percentage error of 2.41, 3.18, and 2.76%, respectively, with a coefficient of determination reaching 0.968, and its overall performance is superior to that of individual models and simple weighted fusion models. The study indicates that a multi-model fusion method based on model diversity and a leakage-free training mechanism can improve the accuracy, robustness, and generalization ability of complex data prediction.

Key words: multi-model fusion; data prediction; stacking; long short-term memory network; XGBoost

参考文献 References

[1]张建勋,胡少杰,芦丽旭,等.多模型融合的时间序列数据预测方法[J].西安邮电大学学报,2025,30(1):115-122.

[2] 周卓辉,杨欢,刘小芳.基于时间模式注意力机制和改进TCN的PM2.5浓度预测方法[J].无线电工程,2024,54(10):2315-2324.

[3] 李浩,卢朝阳,谈翌平,等. 基于MIC-iAFFStacking 集成学习的航空器滑出时间预测[J].交通运输工程与信息学报,2024,22(4):142-153. 

[4] 代业明,周琼.基于改进Bi-LSTM和XGBoost的电力负荷组合预测方法[J].上海理工大学学报,2022,44(2):138-147. 

[5]康文豪,徐天奇,王阳光,等.基于CEEMDAN精细复合多尺度熵和Stacking集成学习的短期风电功率预测[J].水利水电技术(中英文),2022,53(2):163-172.

[6]仲浩帆,黎雅红,朱恩豪.基于模型组合的电力负荷精准预测[J].自动化应用,2024,65(7):223-229.

引用本文

黄欣欣 唐乐红 黄长泉 陈海玥 叶 晨 , 基于多模型融合的数据预测方法研究[J]. 社会发展观察, 2026; 1: (3) : 9-12.