PEER-REVIEW ARTICLE PEER-REVIEWED ARTICLE bioresources.com Wu et al. (2023). “FTIR classification of alfalfa hay,” BioResources 18(3), 5399-5416. 5399 Classification of Alfalfa Hay Based on Infrared Spectroscopy Xiaoqing Wu,a,b Guifang Wu,a,* Bo Wang,a and Jie Li a Alfalfa hay plays a decisive role in the quality and safety of livestock products. Chemical analytical methods for alfalfa hays are laborious, time- consuming, and costly. Therefore, suitable methods are required for rapid and accurate detection of alfalfa hay. This study evaluated the feasibility of infrared spectroscopy (IR) in identifying different alfalfa hays. 105 alfalfa hay samples under three different drying methods were analysed. Results indicated that the full spectra model constructed through standard normal variable transformation (SNV), first-derivative (FD), and second-derivative (SD) preprocessing by BP and SVM had the best performance. The accuracies were all up to 100%. Under the same preprocessing method, the accuracy of BP neural networks was better than that of support vector machine models in most cases. The characteristic wavelength-based SNV-SD-SPA by BP exhibited better performance than the other pretreatment methods, such as: SNV-SPA, SNV-FD-SPA, and SNV-GA, etc. The classification accuracy of moldy-dried alfalfa, sun-dried alfalfa, and shade-dried alfalfa in the training set were 100%, 100%, and 99.5%, respectively, and the accuracy of the prediction set reached 100%, 97.6%, and 97.4%, respectively. Thus, a better theoretical basis was obtained for the grading and online monitoring of alfalfa hay. DOI: 10.15376/biores.18.3.5399-5416 Keywords: Alfalfa hay; Infrared spectroscopy; Machine learning; Classification Contact information: a: College of Mechanical & Electrical Engineering, Inner Mongolia Agricultural University, Hohhot, 010018, P.R. China; b: College of Physics and Electronic Information, Inner Mongolia Normal University, Hohhot, 010022, P.R. China; *Corresponding author: wgfsara@imau.edu.cn INTRODUCTION Alfalfa hay is an important feed for dairy cows and plays an important role in the healthy and stable development of the dairy market (Darabighane et al. 2020; Lorenzo et al. 2020). Dried alfalfa needs to be compressed and processed into a certain size bale for storage and transportation (Cheng et al. 2018; Vanzant et al. 1990). To ensure alfalfa nutrition, bundling is carried out according to a certain water content (Han et al. 2004; Lim et al. 2020). Improper antimildew measures are conducive to the proliferation of microorganisms and cause alfalfa mildew (Wang et al. 1996). The nutrient content of alfalfa after mildew infestation is destroyed, and its feeding value is low, which can cause livestock poisoning and affect the milk product quality (Coblentz et al. 1996). Therefore, if alfalfa mildew can be quickly identified during drying or storage, the loss can be effectively reduced. Traditional methods of chemical detection of mold generally have the characteristics of cumbersome operation, long detection period, and high cost (Gfrerer et al. 2004). Infrared spectroscopy technology is a chemical analysis method that can detect different absorbance frequencies of specific molecules in substances and is fast and PEER-REVIEWED ARTICLE bioresources.com Wu et al. (2023). “FTIR classification of alfalfa hay,” BioResources 18(3), 5399-5416. 5400 nondestructive (Xiong et al. 2016; Zhou et al. 2022). Because different chemical components contain different chemical groups, corresponding to different group frequencies, the positions of the generated characteristic absorbance peaks are also different, and moreover, for the same chemical composition, the intensity of the characteristic absorbance peaks reflected by the different content is not the same (Hell et al. 2016). Therefore, for both quantitative and qualitative analyses of substances, infrared spectroscopy can be utilized. Traditional mid-infrared spectroscopic analysis requires the production of potassium bromide tablets for solid samples. Attenuated total reflection (ATR) technology obtains the information through the reflection signal of the sample surface (Undugodage et al. 2018). It has the characteristics of high sensitivity, clearly characteristic bands, simple operation, and there is no need for sample preparation (Kuronuma et al. 2020). However, the intensity of the overall signal is hard to control when using ATR plate methods, since it depends on the smoothness and pressure of pressing the specimen onto the plate. In recent years, infrared spectroscopy techniques have been widely used in food, pharmaceutical spetrochemicals, tea, wood, feed, and other fields (Mohebby 2010; Wallén et al. 2018; Zapata et al. 2021). There are also some research reports on the detection of food mildew by infrared spectroscopy. Shen and Huang established an online corn mold detection system using spectra and image information fusion technology, by collecting the spectra and image information of corn samples stored on days 6, 9, 12, and 15 and establishing the discriminant linear analysis model, an overall recognition rate of 91.1% was obtained for different degrees of mildew (Shen and Huang 2019). Chu et al. (2014) used near-infrared spectroscopy technology to detect corn kernels with different degrees of mildew. They used principal component analysis to reduce the dimensionality of the spectral data and established a model with FDA (Fisher discriminant analysis), which had a classification accuracy of 91.4%. At present, most studies use spectroscopy and machine vision techniques to detect mildew in food, and there are few reports on the use of infrared spectroscopy to detect mildew in alfalfa hay. Infrared spectral data has the characteristic of high correlation between two adjacent spectra data and high dimensionality (Tanaka et al. 2011). Using full-spectra data to build a model will increase the computing time, and the recognition and classification results may not be ideal. With the development of computer science and artificial intelligence, more machine learning algorithms have been developed and applied to information mining of infrared spectra. Machine learning is a field of study that automatically detects patterns in data from a given database of knowledge and then uses the detected patterns to predict unknown data. Therefore, infrared spectroscopy combined with machine learning may be a potential solution for identifying the quality of alfalfa hay (Kumar et al. 2017). The objectives of this study were as follows: (1) to acquire spectra of alfalfa hay, (2) to determine the optimal wavelength using successive projection algorithm (SPA) and genetic algorithm (GA), (3) to construct a classification model by using the full spectra and optimal wavelengths, and (4) to use neural network and SVM procedures to classify the extracted features. PEER-REVIEWED ARTICLE bioresources.com Wu et al. (2023). “FTIR classification of alfalfa hay,” BioResources 18(3), 5399-5416. 5401 EXPERIMENTAL Preparation of Experimental Samples Samples used for this study were collected from an experimental field of Inner Mongolia Agricultural University in 2019. They were split into three categories of dry alfalfa: alfalfa dried in the shade, alfalfa naturally dried in the sun, and moldy naturally dried in the sun. The alfalfa moisture content was approximately 15% to 20%. Three different types of dry alfalfa were first ground into powder using an electric high-speed pulverizer, which was followed by a 1-mm mesh sieve to remove impurities. Finally, 5 grams of the powdered alfalfa were weighed using an electronic scale into a 50-mL test tube and covered with plastic wrap for storage. Thirty-five samples were prepared for each type of dried alfalfa, and all 105 samples were prepared. Infrared Spectral Acquisition Infrared spectra were recorded using an attenuated total reflectance sampling accessory (PerkinElmer, Boston, MA, USA) and PE series software. Reflectance data were recorded over the wavenumber range of 400 to 4000 cm-1 with 64 scans per spectra and a spectral resolution of 4 cm-1. The acquisition time for a single spectra was 66 s. Background spectra were collected with no samples present on the crystal, and under the same experimental conditions. To assess repeatability and identify any problems caused by the sample's finite particle size, three spectra for each sample were gathered. For the statistical analysis, the average of these three spectra for each sample was then used. The files were exported as comma separated value (csv) files and imported into the MATLAB software (Mathworks, 2020b, Natick, MA, USA) for analysis. The first and last noise of the spectral data was relatively large, and finally 600 to 2000 cm-1 were retained. There were 701 variables in each spectra for preprocessing and modelling in the study. Methods Pretreatment of the spectral data Pretreatment of the averaged spectra was required to eliminate mechanical noise and baseline drift. Pretreatment methods include standard normal variable (SNV), MSC (multiplicative scatter correction), first derivatives (FD), second derivatives (SD), and Savitzky-Golay convolution smoothing (SG), and so on. In order to eliminate strength differences between different samples and analyze data, all data were normalized before preprocessing. The standard normal variable transformation is primarily used for the surface scattering influence and light intensity changes on the spectra, and multivariate scattering correction is used to eliminate the influence of particle size and scattering caused by particle inhomogeneity (Kamruzzaman et al. 2016). A derivative operation is used to eliminate the shift of the baseline. The SG smoothing improves the smoothness of the spectra and reduces the interference of noise (Rahman et al. 2016). In this study, spectral data preprocessing was performed using Unscrambler 10.1 (Camo Software, Oslo, Norway). Characteristic Wavelength Selection Infrared spectral data contain hundreds of continuous wavelengths, which are redundant and multicollinear. Eliminating redundant wavelengths and selecting optimal variables not only can simplify the modelling process and reduce costs and running time, but also, they can improve the performance of the model. In this study, PEER-REVIEWED ARTICLE bioresources.com Wu et al. (2023). “FTIR classification of alfalfa hay,” BioResources 18(3), 5399-5416. 5402 the uninformative variable elimination (UVE)-SPA and GA methods were selected to extract the optimal wavelengths in MATLAB (Version 2020a, MathWorks, Natick, MA, USA). In the UVE-SPA method, UVE can remove a lot of invalid information. Variable modelling based on UVE selection can avoid model overfitting and improve its predictive ability. The SPA mainly solves the problem of collinearity, and it is used to select the wavenumber with the lowest redundant information and obtain the useful variable with the least collinearity (Mário et al. 2001). SPA has been widely used in the selection of spectral characteristic variables. The basic principle of the SPA is to simply project a set of wavelength subsets into the vector space and select the wavelength subset with the least redundancy. The algorithm steps are described below, assuming that the first wavelength k(0) and N are given. Step1: Before the first iteration (n=1),let 𝑥𝑗 = 𝑗 𝑡ℎ 𝑐𝑜𝑙𝑢𝑚𝑛 𝑜𝑓 𝑋𝑐𝑎𝑙; 𝑗 = 1, … , 𝐽. Step2: Let S be the set of wavelength which have not been selected yet. S = {𝑗 𝑠𝑢𝑐ℎ 𝑡ℎ𝑎𝑡 1 j J and j {k(0), … , k(n − 1)}}. Step3: Calculate the projection of 𝑥𝑗 on the subspace orthogonal to 𝑥𝑘(𝑛−1)as P𝑥𝑗 = 𝑥𝑗 − (𝑥𝑗 𝑇𝑥𝑘(𝑛−1))𝑥𝑘(𝑛−1)(𝑥𝑘(𝑛−1) 𝑇 𝑥𝑘(𝑛−1)) −1 for all 𝑗 ∈ 𝑆, where P is the projection operator. Step4: Let 𝑘(𝑛) = arg (max‖P𝑥𝑗‖, 𝑗 ∈ 𝑆). Step5: Let 𝑥𝑗 = P𝑥𝑗, 𝑗 ∈ 𝑆. Step6: Let 𝑛 = n + 1, if n