Abstract:The heavy reliance on imports for China’s feed protein resources necessitates the exploitation of unconventional protein sources to fundamentally address the domestic supply shortage. However, the bioconversion of non-grain biomass into microbial feed protein is characterized by complex multi-parameter coupling. Traditional optimization methods often struggle to decipher the intricate non-linear relationships among variables, resulting in limited efficiency. To overcome these bottlenecks, this study aims to develop high-fidelity predictive models based on machine learning for the precise and efficient optimization of fermentation conditions. To address these shortcomings, we employed machine learning for process optimization. On the basis of an existing dataset from solid-state fermentation of non-grain biomass for feed protein production, we used 13 fermentation parameters, including time, temperature, pH, and nitrogen source input, as input variables, with the true protein yield as the response variable. We used both hold-out and leave-one-out cross-validation methods to compare the predictive performance of three linear and three nonlinear models. Our results demonstrated that under leave-one-out cross-validation, nonlinear models outperformed linear models in predicting the true protein yield. The CatBoost model achieved the best performance, with a coefficient of determination (R2) of 0.84 and a root mean square error (RMSE) of 0.98. Using the trained CatBoost model, we evaluated 8 225 290 fermentation condition combinations by setting stepwise variations for eight key process variables. The model predicted maximum true protein yields for three specific substrate-microbial consortium combinations that exceeded the highest experimental values recorded in our original dataset. We then conducted experimental validation under the model-optimized conditions for the combination of ammonia-pretreated wheat straw and an acclimated microbial consortium. This experiment achieved a true protein yield of 16.51%, which represented a 18.43% improvement over that of the pre-optimization level. Furthermore, the optimized process reduced fermentation time by 108 h and nitrogen source input by 40%, indicating significant potential for cost reduction. This study demonstrates that machine learning can effectively improve the optimization efficiency of solid-state fermentation for feed protein production, thereby laying solid theoretical and practical foundations for developing non-grain biomass resources and alleviating the feed protein shortage in China.