Available online at www.HighTechJournal.org HighTech and Innovation Journal Vol. 5, No. 4, December, 2024 885 ISSN: 2723-9535 A Method for Assessing Urban Industrial Ecological Efficiency Using SBM-GML Model with Tax Reduction Yuchen Guo 1, Jianwei Guo 2* 1 College of Social Science, University of Glasgow, Glasgow G12 8QQ, United Kingdom. 2 School of Traffic and Transportation, Beijing Jiaotong University, Beijing 100044, China. Received 13 August 2024; Revised 17 November 2024; Accepted 22 November 2024; Published 01 December 2024 Abstract The industrial development of cities promotes social and economic development, but it also affects cities' ecological environments. To balance the relationship between the two, the country introduces corresponding tax reduction policies as an effective means of regulation. Therefore, to explore tax and fee reduction policies' specific impact on urban industrial ecological efficiency, the proposed text clustering model was first used in this experiment to cluster the tax and fee reduction policies issued by the government. Subsequently, the Slack-Based Measure-Global Malmquist Lunberger was constructed to measure urban industrial development's ecological efficiency. These experiments confirmed that policy text clustering models had different clustering accuracy on different datasets, with clustering accuracy reaching up to 80.95%, 87.13%, and 94.08% at iterations of 200, 500, and 1000. The regression coefficients for the main variables obtained from the clustering policy, including overall tax reduction and fee reduction, circulation tax reduction, income tax reduction, social expense reduction, and technological innovation tax reduction, were 0.117, 0.105, 0.269, 0.112, and 0.115, respectively. This indicated that these tax and fee reduction measures affected industrial ecological efficiency positively. Therefore, the proposed method can effectively cluster policy texts and measure the industrial ecological efficiency of cities, which has practical feasibility. This provides an effective path for promoting industry and the ecological environment's balanced development. Keywords: Tax Reduction Policy; LDA Text Clustering; SBM-GML; Industrial Ecological Efficiency; Measure. 1. Introduction As an important pillar supporting national economic and social development, the vigorous development of industries ensures people's livelihoods and stable economic growth [1]. However, China's industrial development always relies on a high-consumption and high-pollution industrial structure, which has adverse effects on environmental resources. Therefore, the country has begun to advocate for the development of the green industry and has introduced corresponding policy support to optimize industrial structure and reduce ecological environmental pressure [2]. Tax Reduction Policy (TRP), as an effective means, can promote industry and ecological environment's coordinated development to a certain extent. The concept of industrial Ecological Efficiency (EE) balances industrial economy and ecological environment and attaches great importance to industrial economic output and environmental pollution issues [3]. Meanwhile, the impact of TRP on industrial EE is relatively complex [4]. A suitable TRP can serve as a new path to measure the efficiency of urban industrial ecology. Therefore, effective text clustering of TRP is also an important step in measuring industrial EE [5]. * Corresponding author: 19114023@bjtu.edu.cn http://dx.doi.org/10.28991/HIJ-2024-05-04-02  This is an open access article under the CC-BY license (https://creativecommons.org/licenses/by/4.0/). © Authors retain all copyrights. https://creativecommons.org/licenses/by/4.0/ https://orcid.org/0009-0008-0708-5322 HighTech and Innovation Journal Vol. 5, No. 4, December, 2024 886 Recently, with the developing industry, many researchers have discussed EE. From the perspective of EE of high- tech companies, Vaičiukynas et al. calculated the social EE of high-tech companies when their impact on industrial EE was complex. This could be achieved through the panel data variable Malmquist index of data envelopment analysis. The social EE of high-tech companies was closely related to their financial situation and socio-economic factors, and intensified competition could cause significant fluctuations in social EE. This study could provide a reference value for investigating the correlation patterns and EE of high-tech companies in different industries and around the world [6]. Xu et al. measured the EE of the Yellow River Basin in China using a stochastic frontier analysis model that included time trend variables. This study mainly measured the land-intensive use efficiency and ecological welfare performance of 57 cities over the past decade. Meanwhile, the coupling model and distribution dynamics theory were used in this experiment to analyze the prime movers. The dominant factors of urban EE were social and natural factors, both of which exhibited dual factor enhancement and non-linear enhancement effects [7]. Ke et al. measured urban green EE innovation for SBM. In addition, to explain the impact of economic development on urban green EE, this study fully utilized Hansen threshold regression and mediation effect models. The industrial and energy structures of a city played a significant mediating role in the urban green EE innovation [8]. Gill et al. used a nonlinear autoregressive distributed lag model to test the relationship between urban financial development and its efficiency. The positive and negative impacts of financial development had different impacts on urban EE, and they had an asymmetric relationship. Therefore, cities should adhere to sustainable development and strive to promote the widespread development and practice of green finance [9]. Zhang et al. incorporated EE into the performance evaluation of urban economic transformation to accurately reflect economic transformation. This study used a combined econometric model to conduct a coupled analysis of urban economic transformation mechanisms, and seven typical coal cities were treated as case objects. The transformation effect indicated that 7 cities initially achieved the transformation of economic growth drivers. The quality of transformation indicated a significant improvement in the EE of these cities [10]. The Latent Dirichlet Allocation (LDA) topic model is a commonly used text clustering method, which has research achievements in various scientific fields. B. Yin et al. analyzed abstract texts from the perspective of literature analysis using the LDA topic modeling method. Through LDA, this study identified 7 clear themes. According to the analysis of the development trend of the theme, there was a significant shift in the focus of the theme research, which verified the effectiveness of LDA. This study provided useful reference and application value for research related to blended learning [11]. Shao et al. improved the traditional LDA by introducing an adaptive iterative method to determine the parameter searching convergence. This model was applied to the language classification of news corpora in the metallurgical field and Chinese news corpora. This improved LDA significantly improved classification accuracy compared to traditional LDA and effectively reduced iterations [12]. From the perspective of online courses, Nanda et al. used LDA to determine each course topic to improve the course experience for learners. This study also identified prominent themes in each learner's answer to each question through LDA. The quality of course content, course evaluation and feedback, interaction with teachers, and accessibility of learning materials could affect the learning experience of learners. Thus, the effectiveness of LDA was validated [13]. Weisser et al. used LDA to cluster and mine topic models of short and sparse texts in social media. Taking short, sparse text as an example, this study used LDA to filter out keywords closely related to the topic, thereby verifying the actual effectiveness of LDA. LDA performed well in generating more precise themes, verifying its effectiveness [14]. Xie et al. used LDA for theme modeling and public sentiment analysis from the perspective of public response to crises. The study first collected a large number of Weibo posts using web crawlers and then analyzed the data using LDA text mining technology. Encouraging each other spiritually was significant for the public when facing crisis events. This study indirectly reflected the practical utility of LDA topic modeling [15]. In summary, the concept of EE is widely discussed and has yielded fruitful research results. Meanwhile, the LDA topic clustering model has many applications in text content analysis and clustering. However, existing research methods have many limitations. In terms of EE measurement, many studies are still based on static models, failing to effectively capture the dynamic impact of policy changes on EE. For example, when analyzing the impact of tax reduction policies on industrial EE, existing methods often fail to consider the long-term effects and changes in different time periods after the implementation of the policies. In addition, when evaluating urban industrial EE, the existing studies mainly focus on financial indicators or environmental pollution indicators and lack diversified evaluation dimensions. In terms of LDA topic models, the traditional LDA model highly relies on word frequency analysis in the process of topic modeling, which may lead to insufficient accuracy of topic extraction in the face of complex policy texts. The clustering effect of the LDA model is often affected by the sparsity and ambiguity of the corpus, which makes it difficult to accurately reflect the policy intent and the substance of the content of the final extracted topic. Therefore, this study solves the dynamic problem of EE assessment by constructing a super-efficiency SBM-GML model. This model can comprehensively capture the changes after the implementation of the policy, provide time series analysis of industrial EE, and help understand the actual impact of tax reduction policies at different stages. To overcome the single problem of evaluation indicators, the study combined with a multidimensional index system to comprehensively consider all aspects of EE, including resource input, environmental impact, and economic benefits, so as to improve the comprehensiveness and scientificity of evaluation. Through the introduction of the PC-TFE-LDA method, the analysis ability of policy text was strengthened to solve the shortage of the traditional LDA model's dependence on word HighTech and Innovation Journal Vol. 5, No. 4, December, 2024 887 frequency. The integration of co-occurrence analysis of policy words and topic feature extraction could effectively improve the accuracy of policy text clustering so as to better understand the potential impact of tax reduction policies on industrial EE. This study's innovation lies in (1) a PC-TFE-LDA policy text clustering method is proposed to explore TRP related to industrial EE. (2) In response to the excessive reliance on word frequency analysis and low clustering accuracy in traditional LDA, methods such as policy word co-occurrence, thematic feature word sets, and similarity measurement are introduced. (3) Based on SBM-GML, a super-efficient SBM-GML industrial EE measurement model is introduced to improve the discrimination between decision-making units. 2. Material and Methods First, this study first constructs a TRP text clustering method for PC-TFE-LDA and then constructs an industrial EE measurement model for SBM-GML. Firstly, this study optimizes traditional LDA to improve its overreliance on word frequency analysis and low clustering accuracy. Because of identifying the main themes in the policy text, a super- efficient SBM-GML is constructed to measure the efficiency of urban industrial ecology. 2.1. Text Clustering of Tax Reduction Policies Based on Improved LDA TRP, as a combination strategy continuously introduced by the current government, covers various preferential measures and affects the efficiency of urban industrial ecology. Therefore, text mining and refinement analysis of TRP are crucial [16, 17]. Firstly, LDA is used to systematically extract TRP elements related to industrial EE, laying the foundation for measuring urban industrial EE. This can reveal the dynamic correlation between tax and fee reduction measures and industrial ecological development and enhance the accuracy of policy effectiveness evaluation. LDA is feasible in extracting TRP themes, as it uses statistical methods to extract themes closely related to industrial EE from massive textual data. This helps to analyze the internal structure and evolutionary trends of policy texts. However, this model overly relies on word frequency analysis, making it difficult to deeply understand the underlying semantics and intricate policy logic of policy texts. The accuracy and relevance of its information extraction may be affected to a certain extent, and the clustering accuracy is not high. This is particularly evident when dealing with policy materials that are rich in professional terminology and highly dependent on contextual contexts [18, 19]. In view of this, a PC- TFE-LDA method is proposed. Figure 1 shows the implementation framework of this method. Theme feature words Start Text word bag Feature processingWeight calculation and similarity measurement Policy polarity labeling End Policy co-occurrence word bag Sampling LDA clustering to obtain topic keywords Text preprocessing Similarity measurement Theme related words K-means secondary clustering to obtain policy clustering results Figure 1. Implementation framework of PC-TFE-LDA In Figure 1, in the PC-TFE-LDA method, it improves the clustering effect by combining co-occurrence analysis of policy words and topic feature extraction and can more effectively identify TRP topics related to industrial EE. Before the analysis, the original text data is cleaned and pre-processed to remove irrelevant words, stop words, etc., to ensure the quality of the text data. The Term Frequency-Inverse Document Frequency (TF-IDF) technique is used to calculate the importance of words and screen out key feature words related to policy themes. By analyzing the co-occurrence frequency of policy words, a co-occurrence network is constructed to identify the relationship between each word. The processed text data is input into the LDA model for theme modeling, and the theme features are extracted through multiple iterative optimizations. The K-means clustering algorithm is applied to the extracted subject words, and specific HighTech and Innovation Journal Vol. 5, No. 4, December, 2024 888 policy themes are finely divided so as to realize a systematic analysis of policy texts. The effectiveness of this method lies in the fact that PC-TFE-LDA avoids subject ambiguity caused by word frequency and improves the accuracy of policy subject extraction by considering the co-occurrence relationship of words. This method can deal with complex policy texts, especially when dealing with tax policies that contain a large number of industry-specific terms and complex semantics, and show better results. There are three important steps in this model, namely policy word co- occurrence, topic feature word set, and similarity measurement. They are respectively responsible for solving the sparsity of policy content, extracting thematic features, and constructing knowledge structures. Figure 2 shows the graph models of the three. Ψ Np β ST F(ST) adj,adv,v,(v,noun),else F(ST) T1 TF-IDF T1 T2 βΨ Nt SimNum (a) Policy word co- occurrence graph model (b) Topic feature word construction graph model (c) A graph model for constructing a topic related word set for similarity measurement Figure 2. Graph model for policy word co-occurrence, topic feature word set, and similarity measurement Words' co-occurrence model is mainly based on statistical methods and is a commonly used method in text processing, suitable for processing various types of texts [20, 21]. This study introduces a co-occurrence model to address the policy content sparsity. The relative co-occurrence 𝑅(𝑤𝑥 ∣ 𝑤𝑦) of word 𝑤𝑥 to word 𝑤𝑦 is represented by Equation 1. 𝑅(𝑤𝑥 ∣ 𝑤𝑦) = 𝑓(𝑤𝑥,𝑤𝑦) 𝑓(𝑤𝑦) (1) In Equation 1, 𝑓(𝑤𝑥 , 𝑤𝑦) represents the times that words 𝑤𝑥 , 𝑤𝑦 appear together in the same window unit. 𝑓(𝑤𝑦) represents the times that 𝑤𝑦 appears. The co-occurrence of 𝑤𝑥 , 𝑤𝑦 is expressed using Equation 2. 𝑑(𝑤𝑥 , 𝑤𝑦) = [𝑅(𝑤𝑥 ∣ 𝑤𝑦) + 𝑅(𝑤𝑦 ∣ 𝑤𝑥)]/2 (2) In Equation 2, 𝑑(𝑤𝑥 , 𝑤𝑦) represents the co-occurrence degree of word 𝑤𝑦 to word 𝑤𝑥. Topic feature words on the foundation of part of speech are obtained. Assuming 𝑆𝑇 represents a short text word bag, the policy co-occurrence word bag 𝐹(𝑆𝑇) is represented by Equation 3. 𝐹(𝑆𝑇) = 𝑐(∑ 𝑠(𝑎𝑑𝑗)𝑖 1 ) ∪ 𝑐(∑ 𝑠(𝑎𝑑𝑣)𝑘 1 ) ∪ 𝑐(∑ 𝑠(𝑣) 𝑗 1 ) ∪ 𝑐(∑ ∑ 𝑠(𝑣 + 𝑛𝑜𝑢𝑛)ℎ 1 𝑗 1 ) ∪ 𝑐(∑ 𝑠(𝑒𝑙𝑠𝑒)𝑛 1 ) (3) In Equation 3, 𝑎𝑑𝑗, 𝑎𝑑𝑣, 𝑣, 𝑛𝑜𝑢𝑛, and 𝑒𝑙𝑠𝑒 represent adjectives, adverbs, verbs, nouns, and other parts of speech, respectively, with 𝑖, 𝑘, 𝑗, ℎ, and 𝑛 representing the corresponding number. ∑ 𝑠(𝑎𝑑𝑗)𝑖 1 , ∑ 𝑠(𝑎𝑑𝑣)𝑘 1 , ∑ 𝑠(𝑣) 𝑗 1 , ∑ ∑ 𝑠(𝑣 + 𝑛𝑜𝑢𝑛)ℎ 1 𝑗 1 , and ∑ 𝑠(𝑒𝑙𝑠𝑒)𝑛 1 represent the corresponding bags. 𝑐 represents a constraint condition. The knowledge set derived from word bags are divided into feature words and related words based on the relationship between part of speech and other words. The strong correlation between feature words and thematic attributes is a key indicator for distinguishing themes. Related words co-occur with other thematic attributes and lack distinctiveness. The LDA topic model defines topics through the distribution of "text topic" and "topic word". This study introduces the concept of topic feature words to distinguish short text topics [22, 23]. These words are closely related to the theme and frequently co-occur with theme related words. Although different themes often have unique characteristic words, a single word may also be associated with multiple themes. If 𝐴𝑖 is defined as topic 𝑇's 𝑖th feature word, and 𝑤 is a word in a topic feature word set, then the topic feature word can be represented by Equation 4. 𝑠𝑝 − 𝑤𝑤𝑜𝑟𝑑(𝑤, 𝐴𝑖 ∈ 𝑇) = ∑ 𝑑𝑤∈𝐴𝑖,𝑤 ′≠𝑤 (𝑤,𝑤 ′) (4) In Equation 4, 𝑠𝑝 − 𝑤𝑜𝑟𝑑(𝑤, 𝐴𝑖 ∈ 𝑇) stands for theme feature words, 𝑇 stands for theme words, and 𝐴𝑖 stands for word features. 𝑤 and 𝑤′ represent the topic feature and topic associated word sets' words, respectively. 𝑑(𝑤,𝑤′) is obtained by calculating the co-occurrence of 𝑤,𝑤′. When 𝑑(𝑤,𝑤′) ≥ 1 represents that the topic feature words have distinctiveness, they can be selected for inclusion in a topic feature word set. HighTech and Innovation Journal Vol. 5, No. 4, December, 2024 889 Topic related words refer to words that describe close relationships with each topic corresponding to the topic feature words, represented by Equation 5. 𝑟𝑒𝑙𝑎𝑡𝑖𝑜𝑛(𝑤, 𝐵𝑗 ∈ 𝑇) = ∑ 𝑑𝐴𝑗≠𝐴𝑖,𝑤 ′∈𝐴𝑗,𝑤 ′≠𝑤 (𝑤,𝑤 ′) (5) In Equation 5, 𝑟𝑒𝑙𝑎𝑡𝑖𝑜𝑛(𝑤, 𝐵𝑗 ∈ 𝑇) represents the topic related word. 𝐵𝑗 represents the 𝑗th theme related word of the theme 𝑇. In the feature processing process, TF-IDF is used to block document words to obtain a set of part of speech sequences. The similarity problem between documents is transformed into a vector similarity problem. The similarity is expressed using Equation 6. 𝑆𝑖𝑚(𝑎, 𝑏) = 𝑥1𝑥2+𝑦1𝑦2 √𝑥1 2+𝑦1 2√𝑥2 2+𝑦2 2 (6) In Equation 6, 𝑎[𝑥1, 𝑦1], 𝑏[𝑥2, 𝑦2] represent two different vectors. Different feature word sets for the same topic should remove duplicate words to improve feature extraction for the topic. Subsequently, each topic is processed with topic related and feature words to achieve knowledge extraction, and the generated knowledge is input into LDA for the first clustering. The first clustering obtains the Top30 topic feature words. This study further adopts K-means for the second clustering, and its standard degree function is represented by Equation 7. 𝐸 = ∑ ∑ |𝑋∈𝐶𝑛 𝑘 𝑛=1 𝑋 − �̄�|2 (7) In Equation 7, 𝐸 represents the standard degree function. �̄� represents the central theme of cluster 𝐶𝑛. When the maximum iteration 𝑖𝑡𝑒𝑟𝑚𝑎𝑥 is reached, the algorithm terminates. Figure 3 shows the constructed PC-TFE-LDA graph model. α θ z F(ST) T1 T2 xij Ψ ψ δ Np β w Ψ Nv β B Nd No Nt Figure 3. PC-TFE-LDA graph model This study conducts cluster analysis on relevant policy texts based on PC-TFE-LDA to more accurately extract key categories related to taxes and fees. Figure 4 shows the main process of analyzing tax and fee reduction entries, detailing the key steps from data preprocessing to final clustering. A text set of TRP published by national, provincial, and prefecture level cities over the past decade is first collected. Then, PC-TFE-LDA is used for clustering analysis to obtain a word cloud map and ultimately obtain the distribution of word topics. Finally, specific tax reduction measures under the theme words are further refined. 2.2. Construction of an Industrial EE Measurement Model for SBM-GML Based on Tax Reduction Policies Industrial EE emphasizes improving industrial economic benefits while optimizing resource utilization and minimizing environmental impacts. TRP supports economic transformation and modernization by reducing the tax and administrative burden on enterprises. Therefore, TRP is closely related to the goal of improving industrial EE [24, 25]. After using PC-TFE-LDA for text analysis and screening of relevant policy documents, key policy areas related to the industrial EE impact can be identified. This is particularly reflected in aspects such as circulation tax, income tax, and social security fees. From a structured perspective, Figure 5 shows how TRP affects industrial EE in different tax categories and costs after screening with high-frequency words. HighTech and Innovation Journal Vol. 5, No. 4, December, 2024 890 Website of the Ministry of Finance The website of the State Administration of Taxation Provincial Government Website Provincial and municipal government websites Collect policy texts PC-TFE-LDA model Word frequency calculation word cloud map Topic distribution of words Social security Turnover tax Personal Enterprise Taxable objects Theme analysis Value added tax Corporate income tax Social insurance expenses Business fees R&D expenses Figure 4. The main process of analyzing tax and fee reduction entries Business fees R&D expenses Turnover tax Duty- free Business tax Corporate income tax Return processing Value added tax Business tax and surcharges Retained tax refund Tax deferment and fee deferment Tax incentives Industrial structure Difficulty in financing for small businesses Social insurance expenses Manufac turing enterprises Replacing business tax with value- added tax Personal income tax Release tube clothing Tax refund review Technological innovation Direct Express Enjoyment Vehicle purchase tax Figure 5. Structured perspective on screening high-frequency policy words The model method is a commonly used means for measuring industrial EE. On the foundation of input-output theory, it quantitatively analyzes the input and output of factors through mathematical models and evaluates industrial EE in experiments objectively and accurately. This method requires strict data requirements, but its construction and parameter settings are complex [26]. SBM takes into account inefficient relaxation variables. The Greenness, Productivity, and Health (GML) model compares and analyzes how technological progress affects EE across time. Both SBM and GML can reflect the dynamic changes in EE in detail, demonstrating high applicability in measuring EE. SBM directly introduces relaxation variables into the objective function, solving the relaxation of input-output variables. The implementation process of TRP and its impact often have continuity and dynamic change in time. The traditional efficiency measurement methods are often limited to static analysis and fail to fully consider the timeliness and phased HighTech and Innovation Journal Vol. 5, No. 4, December, 2024 891 effects of policy implementation. The introduction of the SBM-GML model has strengthened the ability to continuously monitor changes in technological progress and economic efficiency, enabling the research to capture the specific impact on EE at various periods during the implementation of tax reduction policies. This dynamic assessment method can accurately reflect the immediate and long-term effects of policies, providing a valuable decision-making basis for enterprises and policy makers. The SBM-GML model combines changes in productivity, technical efficiency, and external environmental impact, and can evaluate industrial EE from multiple dimensions. In the context of the TRP, enterprises not only want to improve their profitability through tax relief but also want to make progress on sustainable development. Therefore, through the adoption of the SBM-GML model, the contribution of different factors to the overall EE can be deeply analyzed, helping policy makers to identify the key factors to improve EE. SBM is represented by Equation 8. 𝑚𝑖𝑛 𝑝 = 1− 1 𝑚 ∑ 𝑠𝑖 −𝑚 𝑖=1 /𝑥𝑖𝑘 1+ 1 𝑞1+𝑞2 (∑ 𝑠𝑟 +𝑞1 𝑟=1 /𝑦𝑟𝑘+∑ 𝑠𝑤 𝑏−𝑞2 𝑤=1 /𝑏𝑤𝑘) 𝑠. 𝑡. { 𝑥𝑘 = 𝑋𝜆 + 𝑠 − 𝑦𝑘 = 𝑌𝜆 − 𝑠 + 𝑏𝑘 = 𝐵𝜆 + 𝑠𝑏− 𝜆, 𝑠−, 𝑠+ ⩾ 0 (8) In Equation 8, 𝑛 refers to the quantity of decision-making units in a production system. 𝑚 refers to the production input in each decision-making unit. 𝑞1 represents the expected output. 𝑞2 represents the type of unexpected output. 𝑋, 𝑌, and 𝐵 represent input vectors, expected output vectors, and unexpected output vectors, respectively. 𝑠−, 𝑠−, and 𝑠𝑏− represent input, expected output, and unexpected output’s slack variables, respectively. 𝜆 represents linear programming’s weight vector. 𝑝 ∈ [0,1] refers to an objective function, which is the efficiency value. If 𝑝 ∈ [0,1], the decision-making unit is effective. If 𝑝 < 1, there is an efficiency loss in the decision-making unit. 𝑋, 𝑋, 𝐵 are represented by Equation 9. { 𝑋 = [𝑥1, … , 𝑥𝑛] ∈ 𝑅 𝑚×𝑛 𝑌 = [𝑦1, … , 𝑦𝑛] ∈ 𝑅 𝑞1×𝑛 𝐵 = [𝑏1, … , 𝑏𝑛] ∈ 𝑅 𝑞2×𝑛 (9) According to Equation 9, the production set can be represented as Equation 10. 𝑃 = {(𝑥, 𝑦, 𝑏) ∣ 𝑥 ⩾ 𝑋𝜆, 𝑦 ⩽ 𝑌𝜆, 𝑏 ⩾ 𝐵𝜆} (10) In Equation 10, 𝑃 represents the production set. This study adopts super-efficient SBM instead of traditional SBM to improve the discrimination between decision-making units. This model allows efficiency values to exceed 1, making the performance of the decision-making unit more prominent compared to other units. By introducing the consideration of unexpected outputs, this model not only measures traditional efficiency but also evaluates the impact of adverse ecological outputs, making it more suitable for EE analysis. The super-efficient SBM is represented by Equation 11. 𝑚𝑖𝑛 𝑝 = 1+ 1 𝑚 ∑ 𝑠𝑖 −𝑚 𝑖=1 /𝑥𝑖𝑘 1− 1 𝑞1+𝑞2 (∑ 𝑠𝑟 +𝑞1 𝑟=1 /𝑦𝑟𝑘+∑ 𝑠𝑤 𝑏−𝑞2 𝑤=1 /𝑏𝑤𝑘) 𝑠. 𝑡. { 1 − 1 𝑞1+𝑞2 (∑ 𝑠𝑟 +𝑞1 𝑟=1 /𝑦𝑟𝑘 +∑ 𝑠𝑤 𝑏−𝑞2 𝑤=1 /𝑏𝑤𝑘) > 0 𝜆, 𝑠−, 𝑠+ ⩾ 0 1 − 1 𝑞1+𝑞2 (∑ 𝑠𝑟 +𝑞1 𝑟=1 /𝑦𝑟𝑘 +∑ 𝑠𝑤 𝑏−𝑞2 𝑤=1 /𝑏𝑤𝑘) > 0 (11) The GML index analyzes the dynamic development of urban industrial EE by calculating the productivity changes of decision-making units at different periods [27]. This index is divided into Greenness Technology Change (GTC) and Greenness Efficiency Change (GEC). GTC mainly measures technological innovation and progress in industrial production, while GEC focuses on the performance of efficiency improvement. The combination of these two indices can effectively describe the overall trend of EE, and can also be used to analyze the potential mediating impact of TRP on industrial EE. GML is represented by Equation 12. 𝐺𝑀𝐿𝑡 𝑡+1 = 1+𝑠𝑉 𝐺(𝑥𝑡,𝑦𝑡,𝑏𝑡,𝑔𝑥,𝑔𝑦,𝑔𝑏) 1+𝑠𝑉 𝐺(𝑥𝑡+1,𝑦𝑡+1,𝑏𝑡+1,𝑔𝑥,𝑔𝑦,𝑔𝑏) = 𝐺𝐸𝐶𝑡 𝑡+1 ∗ 𝐺𝑇𝐶𝑡 𝑡+1 (12) In Equation 12, 𝑠𝑉 𝐺 represents the directional distance function of SBM. 𝑡 represents a specific period. 𝑔𝑥, 𝑔𝑦, and 𝑔𝑏 refer to the direction vectors for reducing input, increasing "good output", and decreasing "bad output", respectively. GEC is represented by Equation 13. 𝐺𝐸𝐶𝑡 𝑡+1 = 1+𝑠𝑉 𝑡 (𝑥𝑡,𝑦𝑡,𝑏𝑡,𝑔𝑥,𝑔𝑦,𝑔𝑏) 1+𝑠𝑉 𝑡+1(𝑥𝑡+1,𝑦𝑡+1,𝑏𝑡+1,𝑔𝑥,𝑔𝑦 ,𝑔𝑏) (13) The function of GTC is represented by Equation 14. HighTech and Innovation Journal Vol. 5, No. 4, December, 2024 892 𝐺𝑇𝐶𝑡 𝑡+1 = {[1+𝑠𝑉 𝐺(𝑥𝑡,𝑦𝑡,𝑏𝑡,𝑔𝑥,𝑔𝑦 ,𝑔𝑏)]/[1+𝑠𝑉 𝑡 (𝑥𝑡,𝑦𝑡,𝑏𝑡,𝑔𝑥,𝑔𝑦,𝑔𝑏)]} {[1+𝑠𝑉 𝐺(𝑥𝑡+1,𝑦𝑡+1,𝑏𝑡+1,𝑔𝑥,𝑔𝑦,𝑔𝑏)]/[1+𝑠𝑉 𝑡+1(𝑥𝑡+1,𝑦𝑡+1,𝑏𝑡+1,𝑔𝑥,𝑔𝑦,𝑔𝑏)]} (14) GML evaluates the dynamic industrial EE changes by comparing index values over continuous periods. Specifically, if the GML index value > 1, this showcases an improvement in industrial EE compared to the previous period during the inspection period, while conversely, it indicates a decrease in efficiency [28, 29]. Evaluating a certain region's industrial EE requires comprehensive consideration of various aspects. A single indicator is insufficient to reflect the dynamic EE changes and has significant limitations. Combining multiple factors for comprehensive evaluation can achieve a multi-dimensional evaluation of the ecological benefits of regional industries [30]. Therefore, in Figure 6, the basis for constructing the evaluation index system of urban industrial EE is summarized using different principles such as scientificity, systematicity, and practicality. The selected evaluation indicators need to be representative and targeted, able to summarize the characteristics of the research object at all levels. The method used needs to be supported by scientific theories and eliminate external human factors to objectively describe the research object. The raw data required for specific evaluation indicators should be easily obtainable, and the calculation method should comply with conventional methods Systematic Scientificity Practicality Principles for constructing an evaluation index system for industrial ecological efficiency: ① The Concept and Connotation of Efficiency in the Context of Tax Reduction , Fee Reduction, and Green Development. ② Sustainable development of economy, resources and environment, i.e. the impact on various systems. ③ Considering the input framework in the Cobb Douglas production function. ④ Overall operational efficiency, including structural indicators, functional status indicators, and process change characteristic indicators. ⑤ The characteristics and actual situation of industrial development in the research area. Figure 6. Basis for constructing urban industrial EE's evaluation index system In summary, this study systematically analyzes the dynamic changes and influencing factors of regional EE by constructing an indicator system. This indicator system is divided into four levels: objectives, systems, criteria, and indicator layers [31]. The target layer is the overall industrial EE. The system layer is subdivided into three subsystems, namely resource input, expected output, and unexpected output, to synthetically reflect the industrial activities EE. The criterion layer measures the functions and outputs of each subsystem, mainly including capital input, production factor input, economic benefits, and environmental pollution. The indicator layer is further refined, mainly including 8 specific evaluation variables. The SBM-GML model considers the balance between undesired outputs (such as environmental pollution) and expected outputs (such as economic gains) in efficiency assessment. The GML part of the model can reflect technological progress and efficiency changes over different time periods, providing support for understanding the long-term impact of tax reduction policies. Therefore, SBM-GML provides a dynamic and comprehensive perspective to evaluate policy effects and is a reasonable choice to measure EE. The SBM-GML model can be applied to many types of data, including incomplete or unbalanced datasets. This makes the model more adaptable in actual operations and can handle data problems that are common in reality. Compared to other models, such as data envelopment analysis models, there is no way to reflect the impact of time changes. The effects of tax reduction often take time to show, so it is important to use models that capture dynamic changes. The SBM-GML model performs well when dealing with unbalanced or incomplete data, while many other models have more stringent data requirements and may not be able to effectively deal with various data problems commonly encountered in practical applications. This constructed evaluation system refers to existing research results and can achieve scientific evaluation of urban industrial EE. Figure 7 shows the overall evaluation system. HighTech and Innovation Journal Vol. 5, No. 4, December, 2024 893 Expected output Urban industrial ecological efficiency Target layer Primary indicators Secondary indicators Input Input of production factors Capital investment Unexpected output Economic benefits Environmental pollution Total consumption (10000 tons of standard coal) Total industrial assets (100 million yuan) Human capital and industrial energy in the secondary industry Industrial added value (100 million yuan) Industrial solid waste generation (10000 tons) Industrial wastewater discharge (10000 tons) Industrial smoke and dust emissions (tons) Industrial sulfur dioxide emissions (tons) Indicator Description Figure 7. Evaluation index system for urban industrial EE In addition, this study uses the TRP clustering results obtained from PC-TFE-LDA as the core explanatory variable to investigate how TRP affects urban industrial EE. Meanwhile, an empirical analysis is conducted on the impact of urban TRP on industrial EE, with industrial EE as a dependent variable. 3. Results and Discussion First, this study validated the clustering performance of PC-TFE-LDA by selecting three policy related datasets: Lexis Nexis, Bloomberg Law, and Data verse, and comparing them with four other popular models. Subsequently, this study used a province as an example to measure urban industrial EE using super-efficient SBM-GML. 3.1. Cluster Effect Analysis of Tax Reduction Policies Based on PC-TFE-LDA To evaluate the proposed PC-TFE-LDA, this study used three datasets suitable for policy cluster evaluation, totaling 5050 entries. These three datasets were Lexis Nexis, Bloomberg Law, and Data verse, respectively. Among them, LexisNexis, as a widely used legal information platform, provides a large number of policy texts and legal provisions, which is suitable for policy analysis and cluster research. The breadth and depth of its content can ensure that the extracted policy topics have legal and regulatory authority. Due to its main legal content, Bloomberg Law can provide the latest and most comprehensive information related to financial and tax policies, which is suitable for analyzing the overall effect of policies in combination with the economic and legal background. Data verse is an academic data archiving platform maintained by Harvard University that supports the storage, sharing, and referencing of research data. By collecting data in various forms, Data verse provides researchers with a wealth of policy-related data that can effectively support in-depth analysis of government policies and social science research. The three datasets contain different types of content, so that different policy texts can be comprehensively analyzed from multiple perspectives, ensuring that the extracted topics have interdisciplinary perspectives and diverse representation; Although these platforms are primarily focused on U.S. legal and policy documents, they can still provide the basis for analyzing and comparing policies in other regions, especially when it comes to global policies or topics of universal applicability. As mainstream legal databases, LexisNexis and Bloomberg Law collect policy texts with high accuracy and authority after strict review. Most of the Data contained in the Data verse has been verified by the academic community, which helps to improve the credibility of the research results. The best policy implementation effect was considered positive data, while the average or poor effect was negative data. To achieve clustering performance testing of PC-TFE-LDA, the experimental environment was Python 3.6 software, with IntelCoreI5-7200U@2.50GHz CPU, 8.00GB memory, and Windows 7 operating system. Figure 8 shows the clustering performance of PC-TFE-LDA on three datasets. Different datasets exhibited their unique optimal topic feature words. In Figure 8 (a), in Lexis Nexis, at iterations of 200, 500, and 1000, when the subject words were set to 15 (K=15), the clustering accuracy reached the highest (81.34%, 85.22%, 92.45%). This indicated that the ideal topic feature words for this dataset were 15 (Top15). In Figure 8 (b), in Bloomberg Law, at iterations of 200, 500, and 1000, the clustering accuracy was highest at K=10 (79.95%, 86.97%, 92.66%). Therefore, the optimal topic feature words were set to 10 (Top10). In Figure 8 (c), in the Data verse, at iterations of 200, 500, and 1000, the clustering accuracy was highest at K=20 (80.95%, 87.13%, 94.08%), indicating that the optimal topic feature words were 20 (Top 20). The results showed that with the increase of the number in iterations, the model could dynamically adjust the text data and identify the theme and its feature words better. At the same time, it reflected the differences in HighTech and Innovation Journal Vol. 5, No. 4, December, 2024 894 information density and topic complexity of different types of texts and promoted the in-depth understanding of data background. Too few feature words make it difficult to clarify the topic, while too many will increase noise, both of which will reduce the clustering effect. 5 10 15 20 25 30 TopK A cc ur ac y (% ) (a) LexisNexis 200 500 1000 5 10 15 20 25 30 A cc ur ac y (% ) (b) Bloomberg Law 10 15 20 25 TopK A cc ur ac y (% ) (c) Dataverse 5 30 0.95 0.90 0.85 0.80 0.75 0.70 TopK 200 500 1000 200 500 1000 0.95 0.90 0.85 0.80 0.75 0.70 0.95 0.90 0.85 0.80 0.75 0.70 Figure 8. The clustering accuracy of PC-TFE-LDA on three datasets Figure 9 shows the fitting performance of PC-TFE-LDA on three policy datasets. From the figure, PC-TFE-LDA all showed a good clustering effect on the three data sets, and the higher the stability of the fitting effect, the stronger the robustness of the algorithm in processing diverse data, which enhanced the trust in the analysis results of policy text and verifies its effectiveness. The main reason is that PC-TFE-LDA obtains the most suitable optimal topic feature words on the Data verse. A moderate number of feature words can help improve clustering performance. If there are too few characteristic words, the theme is not clear. Too many feature words can easily generate noise interference. 0 0.2 0.4 0.6 0.8 1.0 0 E st im at e va lu e Actual value 0 E st im at e va lu e 0.2 0.4 0.6 0.8 1.0 0 E st im at e va lu e 0.2 0.4 0.6 0.8 1.0 1.2 0.2 0.4 0.6 0.8 1.0 1.2 0.2 0.4 0.6 0.8 1.0 0.2 0.4 0.6 0.8 1.0 Actual evaluation results Predicting evaluation results Actual value Actual value Actual evaluation results Predicting evaluation results Actual evaluation results Predicting evaluation results (a) LexisNexis (b) Bloomberg Law (c) Dataverse Figure 9. The fitting effect of PC-TFE-LDA on three policy datasets HighTech and Innovation Journal Vol. 5, No. 4, December, 2024 895 This study compared PC-TFE-LDA with the Joint Sentiment Topic Model (JST), Latent Semantic Model (LSM), Labeled Topic Model (LTM), and Enhanced Latent Dirichlet Allocation (ELDA) on multiple metrics including Precision, Recall, and F-measure. Figure 10 shows the clustering results of positive and negative pole data for five models. In Figure 10 (a), for the positive electrode data, these indicators of PC-TFE-LDA fluctuated around 0.90, all better than JST, LSM, LTM, and ELDA. In Figure 10 (b), for the negative electrode data, these indicators of PC- TFE-LDA also fluctuated around 0.90, all of which were better than other models. Therefore, PC-TFE-LDA performed better than the other four models in the accuracy rate of positive and negative data, recall rate, and F - measure value, indicating that this new method had high effectiveness and strong stability in accurately identifying the subject of policy text. When PC-TFE-LDA was used, the relevant features could be better identified and extracted, especially when the policy text was processed, and the design of the model was more targeted, which could significantly improve the clustering effect. F-measureRecallPrecision (a) Positive polarity calculation results 0.4 0.5 0.6 0.7 0.8 0.9 10 F-measureRecallPrecision (b) Negative polarity calculation results 0.4 0.5 0.6 0.7 0.8 0.9 10 JST LSM LTM ELDA SKP-LDAJST LSM LTM ELDA SKP-LDA V al ue (% ) V al ue (% ) Figure 10. Clustering results of positive and negative pole data for 5 models This study selected the topics quantity K to analyze the consistency and confusion of different topic numbers in Figure 11. Low confusion indicates low uncertainty and good effectiveness, while high consistency reflects strong semantic relevance of words under the theme. By comparing the confusion and consistency under different K values, when K=10, the consistency was high and the confusion gradually stabilized. Therefore, it was determined that there were 10 topics. These 10 topics analyzed by LDA were a set of feature words, each of which served as a focal point for a type of policy. Through semantic analysis and summarization, the key words were identified as "social security, turnover tax, individual, enterprise, taxable object, technological innovation, research and development expenses, industrial structure, financing difficulties for small enterprises, and business tax". This study uses tax reduction and fee reduction, circulation tax reduction, income tax reduction, social expense reduction, and technological innovation tax reduction as core explanatory variables. Industrial EE was regarded as a dependent variable to better examine how TRP affects urban industrial EE. -2.9 -3.3 -3.7 -4.1 -4.5 15141312111098765432 0.326 0.327 0.328 0.329 0.330 C on si st en cy Number of themes C on fu si on le ve l Confusion level Consistency Figure 11. Consistency and confusion results of different topic numbers 3.2. Measurement Effect of Industrial EE Measurement Model Using Tax Reduction Policy and Improved SBM-GML This study took 10 cities from 3 regions in a certain province as the research objects and super efficiency SBM-GML was used to empirically analyze their industrial EE. Firstly, a summary analysis was conducted on the industrial HighTech and Innovation Journal Vol. 5, No. 4, December, 2024 896 ecological development status of the province since the implementation of TRP. Then the industrial EE was calculated and summarized using MAXDEA software from the provincial, municipal, and sub-industry dimensions. Figure 12 shows the trend of industrial EE changes and EE GML in the province since the implementation of TRP. In Figure 12 (a), industries above designated size in this province experienced rapid growth from 2014 to 2016. Since 2017, the industrial growth rate gradually slowed down and turned to medium growth. By 2022, it decreased to 1% due to economic shocks. However, the industrial growth rate rebounded to 7.59% in 2023, indicating that its industrial economic growth gradually became rational and maintained stability. In Figure 12 (b), since the implementation of TRP, the overall industrial EE of the province showed a decline followed by an increase, with an average of over 1.041 and an average annual growth rate of 3.31%. Since 2017, industrial EE has been increasing year by year, showing continuous improvement. Further analysis confirmed that technological progress and efficiency were key factors driving the growth of industrial EE. The average annual growth rate of technological progress was 4.79%, the technical efficiency was 3.88%, and the average annual values were 1.139 and 1.052, respectively. This trend indicated that technological innovation and efficiency improvement had simultaneously promoted the continuous optimization of industrial EE. Figure 12 shows that in the initial stage of policy implementation, EE might be affected by the external economic environment, and then gradually recovered, indicating that the implementation of TRP has played a positive role in promoting the development of enterprises. At the same time, since the implementation of the policy, with the improvement of technological progress and technical efficiency, the average industrial EE has remained at a high level, reflecting the success of the TRP in promoting technological innovation. 5 15 20 25 0 10 G ro w th r at e ( % ) Year 30 Industrial output Total industrial output value (a) The trend of growth rate changes in industries above designated size in the province -5 -10 0.92 0.96 0.98 1.00 0.90 0.94 Va lu e Year 1.10 (b) The GML index and decomposition value of industrial efficiency in the province 1.20 1.30 GTECH GEFFCH Efficiency value Figure 12. The trend of industrial EE change in the province and the index value of EE GML Figure 13 shows the comparison results of the average EE and the trend of EE changes in the province and its regions. In Figure 13 (a), Zone B exceeded the provincial average level, with an average industrial EE of 1.115, which was 7.11% higher than the provincial average level. The industrial EE of Zones A and C were 1.037 and 0.961, respectively, which were 0.08% and 7.40% lower than the provincial average level. Compared to Zone B, Zones A and C had lower industrial EE. The results reflected the uneven impact of policy implementation in different regions. This result suggested that policy makers should consider regional differences when making tax reduction plans to better realize the optimal allocation of resources. In Figure 13 (b), the industrial EE of the province showed an overall fluctuating growth since the implementation of TRP. From 0.972 in 2014 to 1.295 in 2023, although it dropped to the lowest value of 0.861 in 2016, it still showed a significant improvement in 2022. The annual growth rate of EE in Zone B was 2.85%, indicating that the growth rate of EE in this area was at a relatively fast level. The industrial EE of Zone A showed a steady increase after a slight initial decline, and the overall efficiency was higher than that of other regions. The industrial EE of Zone A gradually increased from 0.982 in 2014 to 1.121 in 2023. Although slightly inferior to Zone B in the initial stage, it performed better in the later stage. The industrial EE of Zone C started at a relatively low level, reaching 0.873 in 2014 and increasing to 1.071 by 2023. Despite its weak foundation, the annual average growth rate was 2.23%, indicating a sustained and stable improvement trend. The results showed that the EE of all districts increased in fluctuation, indicating that the policy may encounter obstacles in some periods, and the occurrence of these fluctuations should lead to further evaluation and adjustment of the policy. This study used Moran scatter plots to analyze the spatial agglomeration of industrial EE in the province in 2016, 2018, 2020, and 2022 in Figure 14. In 2016, there were three cities with industrial EE located in the first and third quadrants, accounting for 60% of the total sample, showing a strong spatial agglomeration effect. Subsequently, in 2018, cities in the first quadrant (high-high agglomeration) decreased by one, while cities in the third quadrant (low-low agglomeration) increased by one. By 2020, cities in the third quadrant decreased by 2, while cities in the first quadrant remained unchanged. In 2022, cities in the first quadrant increased by 2, while cities in the third quadrant remained unchanged. The variation of the results reflected that the high-high agglomeration trend has been weakened and then HighTech and Innovation Journal Vol. 5, No. 4, December, 2024 897 strengthened. This change showed that under the influence of the policy, regions with low efficiency were gathering in the direction of high efficiency, trying to narrow the development gap between regions. The decrease in the number of low-low agglomerations reflected that the policy implementation has achieved positive results in improving the overall industrial EE, suggesting that the policy has a certain guiding effect. 0.85 0.90 0.95 1.00 1.05 1.10 1.15 E ff ic ie nc y m ea n 1.039 1.038 1.113 0.962 The province CBA (a) Comparison of average ecological efficiency in Shaanxi and sub-regions 1.30 1.25 1.20 1.15 1.10 1.05 1.00 0.95 0.90 0.85 0.80 2014 2015 2016 2017 2018 2019 2020 2021 The province A B C 20232022 (b) Changes in efficiency by region in Shaanxi Province E ff ic ie nc y va lu e Figure 13. Comparison of the mean EE and trend of EE changes in the province and its regions -1 1 -2 0 W e ig h te d a v e ra g e o f e c o lo g ic a l e ff ic ie n c y -1-2 Ecological efficiency 10 2 (a) 2016 High-high agglomerati on (HH) High-low agglomerati on (HL) Low-low agglomerati on (LL) Low-high agglomerati on (LH) -1 1 -2 0 W e ig h te d a v e ra g e o f e c o lo g ic a l e ff ic ie n c y -1-2 Ecological efficiency 10 2 (b) 2018 High-high agglomerati on (HH) High-low agglomerati on (HL) Low-low agglomerati on (LL) Low-high agglomerati on (LH) -1 1 -2 0 W e ig h te d a v e ra g e o f e c o lo g ic a l e ff ic ie n c y -1-2 Ecological efficiency 10 2 (c) 2020 High-high agglomerati on (HH) High-low agglomerati on (HL) Low-low agglomerati on (LL) Low-high agglomerati on (LH) -1 1 -2 0 W e ig h te d a v e ra g e o f e c o lo g ic a l e ff ic ie n c y -1-2 Ecological efficiency 10 2 (d) 2022 High-high agglomerati on (HH) High-low agglomerati on (HL) Low-low agglomerati on (LL) Low-high agglomerati on (LH) Figure 14. Moran scatter plot of industrial EE in the province Finally, this study discussed the specific impacts of various TRPs on industrial EE. Table 1 shows the model regression results for each variable. At a significance level of 5%, there were 5 main variables involved in TRP. They were overall tax reduction and fee reduction, circulation tax reduction, income tax reduction, social expense reduction, and technological innovation tax reduction, all showing statistical significance. The regression coefficients of these variables were 0.117, 0.105, 0.269, 0.112, and 0.115, all of which were positive values. This indicated that these measures had a positive impact on industrial EE. The results emphasized the core position of these tax and fee reduction measures in promoting the enterprise efficiency cycle, indicating that they had certain effectiveness and provided strong evidence for the sustainability of the policy. At the same time, there was a complementary relationship between various tax reduction measures to improve the overall innovation and environmental protection capabilities of enterprises. HighTech and Innovation Journal Vol. 5, No. 4, December, 2024 898 Table 1. Model regression results for each variable Variable Regression coefficient Standard error t P>t 95% confidence interval Overall tax reduction and fee reduction 0.117 0.058 2.011 0.045 0.003 0.232 Circulation tax reduction 0.105 0.024 4.282 0.000 0.155 0.057 Income tax reduction 0.269 0.05 5.373 0.000 0.171 0.369 Social expense tax reduction 0.112 0.052 2.181 0.031 0.218 0.011 Technological innovation tax reduction 0.115 0.030 3.811 0.000 0.057 0.178 The effectiveness of the model was verified in the above experiments. The experimental results showed that the PC- TFE-LDA algorithm could obtain different numbers of subject terms in the datasets LexisNexis, Bloomberg Law, and Data verse, which were K=15, K=10, and K=20, respectively. When the number of iterations was 200, 500, and 1000, the highest clustering accuracy of PC-TFE-LDA on the three datasets was 92.45%, 92.66%, and 94.08%. The comparison results with the JST, LSM, LTM, and ELDA models showed that the accuracy rate, recall rate, and F- measure value of PC-TFE-LDA on positive and negative data sets all fluctuated around 0.90 and were better than the comparison model. The consistency-confusion test results showed that when K=10, the consistency was high and the confusion degree was gradually stable, so the number of topics was determined to be 10. The empirical results of industrial EE based on the super-efficiency SBM-GML model showed that since the implementation of tax and fee reduction policies, the industrial eco-efficiency of the provinces selected in the study has first decreased and then increased, with an average value of 1.041 and an average annual growth rate of 3.31%. The average annual growth rate of technological progress was 4.79%, and the average annual technical efficiency was 3.88%, reaching 1.139 and 1.052, respectively. Compared with previous studies, Lee et al. proposed the method of using green finance to improve EE. The results showed that green finance significantly promoted the improvement of EE, and the higher the EE, the more obvious the improvement effect. The upgrading of industrial structure, optimization of energy structure, enterprises' concern for environmental protection, and the public's concern for the environment were all favorable factors to strengthen the role of green finance in promoting EE [32]. However, the dynamic process of policy implementation was often not deeply analyzed, and the impact of tax reduction policies on EE was not systematically discussed. Through dynamic assessment, the study filled the gap in the short- and long-term impact analysis of previous studies and emphasized the long-term effect of tax reduction policies on technological progress. Guo et al. used the super-efficiency relaxation measurement model to measure the coordination between economic and social development and environmental protection and used the Tobit model to explore the factors affecting the efficiency of alleviating ecological poverty [27]. However, in the relaxation variables of non-expected output, the static analysis was emphasized, and the dynamic effects of time and policy changes on EE were not fully considered. The SBM-GML model introduced time series analysis to better capture the long-term effects and dynamic changes after the implementation of tax reduction policies. 4. Conclusion With the intensification of global concern for sustainable development, how to improve industrial EE and promote green transformation while pursuing economic benefits has become a major issue to be solved urgently. The research adopted the methodology of double innovation. Firstly, the co-occurrence of policy words and topic feature extraction were introduced to construct PC-TFE-LDA to improve the accuracy of policy text analysis. Secondly, the super- efficiency SBM-GML model was constructed to measure the industrial EE dynamically. Through experimental verification, the results emphasized the importance of TRP in promoting technological innovation and fully verified the effectiveness of the proposed method. The super-efficiency SBM-GML model provided a new perspective for dynamic assessment of EE, which made the research have important technical significance in the analysis of empirical results and provided a scientific basis for the formulation and optimization of actual policies. Based on the research results, the following suggestions are put forward to improve the province's industrial EE: continuing to promote circulation tax and income tax reduction to further reduce the burden on enterprises and encouraging more enterprises to increase investment in technological innovation and environmental protection facilities. By encouraging research and development and providing tax incentives, enterprises are supported in technological progress in cleaner production and green technology, thereby improving the overall industrial EE. The limitation of this study is that it failed to consider the influence of the dynamic change index and the long-term effect of the policy implementation. Future research direction will introduce a dynamic change index to analyze its specific role in industrial EE for continuous tracking and evaluation of the long-term effects of policies. 5. Declarations 5.1. Author Contributions Y.G. and J.G. contributed to the design and implementation of the research, to the analysis of the results and to the writing of the manuscript. All authors have read and agreed to the published version of the manuscript. HighTech and Innovation Journal Vol. 5, No. 4, December, 2024 899 5.2. Data Availability Statement The data presented in this study are available on request from the corresponding author. 5.3. Funding The authors received no financial support for the research, authorship, and/or publication of this article. 5.4. Institutional Review Board Statement Not applicable. 5.5. Informed Consent Statement Not applicable. 5.6. Declaration of Competing Interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. 6. References [1] Köppl, A., & Schratzenstaller, M. (2023). Carbon taxation: A review of the empirical literature. Journal of Economic Surveys, 37(4), 1353–1388. doi:10.1111/joes.12531. [2] Guo, B., Wang, Y., Zhou, H., & Hu, F. (2023). Can environmental tax reform promote carbon abatement of resource-based cities? Evidence from a quasi-natural experiment in China. Environmental Science and Pollution Research, 30(55), 117037–117049. doi:10.1007/s11356-022-23669-3. [3] YAN, X., & TU, J. (2021). The spatio-temporal evolution and driving factors of eco-efficiency of resource-based cities in the Yellow River Basin. Journal of Natural Resources, 36(1), 223. doi:10.31497/zrzyxb.20210115. [4] Usman, A., Quan-Lin, L., Abdullah, A. M., Shakib, M., & Wasim, I. (2021). Correction to: Nexus between agro-ecological efficiency and carbon emission transfer: evidence from China. Environmental Science and Pollution Research, 28(32), 44581- 44581. doi:10.1007/s11356-021-14461-w. [5] George, L., & Sumathy, P. (2023). An integrated clustering and BERT framework for improved topic modeling. International Journal of Information Technology (Singapore), 15(4), 2187–2195. doi:10.1007/s41870-023-01268-w. [6] Vaičiukynas, E., Andrijauskienė, M., Danėnas, P., & Benetytė, R. (2023). Socio-eco-efficiency of high-tech companies: a cross- sector and cross-regional study. Environment, Development and Sustainability, 25(11), 12761–12790. doi:10.1007/s10668-022- 02589-9. [7] Xu, W., Xu, Z., & Liu, C. (2021). Coupling analysis of land intensive use efficiency and ecological well-being performance of cities in the Yellow River Basin. Journal of Natural Resources, 36(1), 114. doi:10.31497/zrzyxb.20210108. [8] Ke, H., Dai, S., & Yu, H. (2022). Effect of green innovation efficiency on ecological footprint in 283 Chinese Cities from 2008 to 2018. Environment, Development and Sustainability, 24(2), 2841–2860. doi:10.1007/s10668-021-01556-0. [9] Gill, A. R., Riaz, R., & Ali, M. (2023). The asymmetric impact of financial development on ecological footprint in Pakistan. Environmental Science and Pollution Research, 30(11), 30755–30765. doi:10.1007/s11356-022-24384-9. [10] Zhang, M., Zhang, P., & Li, H. (2021). Characteristics and evaluation methods of economic transformation performance of resource-based cities: An empirical study of Northeast China. Journal of Natural Resources, 36(8), 2051. doi:10.31497/zrzyxb.20210811. [11] Yin, B., & Yuan, C. H. (2022). Detecting latent topics and trends in blended learning using LDA topic modeling. Education and Information Technologies, 27(9), 12689–12712. doi:10.1007/s10639-022-11118-0. [12] Shao, D., Li, C., Huang, C., Xiang, Y., & Yu, Z. (2022). A news classification applied with new text representation based on the improved LDA. Multimedia Tools and Applications, 81(15), 21521–21545. doi:10.1007/s11042-022-12713-6. [13] Nanda, G., A. Douglas, K., R. Waller, D., E. Merzdorf, H., & Goldwasser, D. (2021). Analyzing Large Collections of Open- Ended Feedback from MOOC Learners Using LDA Topic Modeling and Qualitative Analysis. IEEE Transactions on Learning Technologies, 14(2), 146–160. doi:10.1109/TLT.2021.3064798. [14] Weisser, C., Gerloff, C., Thielmann, A., Python, A., Reuter, A., Kneib, T., & Säfken, B. (2023). Pseudo-document simulation for comparing LDA, GSDMM and GPM topic models on short and sparse text using Twitter data. Computational Statistics, 38(2), 647–674. doi:10.1007/s00180-022-01246-z. HighTech and Innovation Journal Vol. 5, No. 4, December, 2024 900 [15] Xie, R., Chu, S. K. W., Chiu, D. K. W., & Wang, Y. (2021). Exploring Public Response to COVID-19 on Weibo with LDA Topic Modeling and Sentiment Analysis. Data and Information Management, 5(1), 86–99. doi:10.2478/dim-2020-0023. [16] Mehta, V., Bawa, S., & Singh, J. (2021). WEClustering: word embeddings based text-clustering technique for large datasets. Complex and Intelligent Systems, 7(6), 3211–3224. doi:10.1007/s40747-021-00512-9. [17] Belwal, R. C., Rai, S., & Gupta, A. (2023). Extractive text summarization using clustering-based topic modeling. Soft Computing, 27(7), 3965–3982. doi:10.1007/s00500-022-07534-6. [18] Saravanan, S., Karthigaivel, R., & Magudeeswaran, V. (2021). A brain tumor image segmentation technique in image processing using ICA-LDA algorithm with ARHE model. Journal of Ambient Intelligence and Humanized Computing, 12(5), 4727–4735. doi:10.1007/s12652-020-01875-6. [19] Sharma, S. K., Vijayakumar, K., Kadam, V. J., & Williamson, S. (2022). Breast cancer prediction from microRNA profiling using random subspace ensemble of LDA classifiers via Bayesian optimization. Multimedia Tools and Applications, 81(29), 41785–41805. doi:10.1007/s11042-021-11653-x. [20] Thielmann, A., Weisser, C., Krenz, A., & Säfken, B. (2023). Unsupervised document classification integrating web scraping, one-class SVM and LDA topic modelling. Journal of Applied Statistics, 50(3), 574–591. doi:10.1080/02664763.2021.1919063. [21] van Ewijk, S., & Hoekman, P. (2021). Emission reduction potentials for academic conference travel. Journal of Industrial Ecology, 25(3), 778–788. doi:10.1111/jiec.13079. [22] Dong, F., & Zheng, L. (2022). The impact of market-incentive environmental regulation on the development of the new energy vehicle industry: a quasi-natural experiment based on China’s dual-credit policy. Environmental Science and Pollution Research, 29(4), 5863–5880. doi:10.1007/s11356-021-16036-1. [23] Tan, L., Wu, X., Guo, J., & Santibanez-Gonzalez, E. D. R. (2022). Assessing the Impacts of COVID-19 on the Industrial Sectors and Economy of China. Risk Analysis, 42(1), 21–39. doi:10.1111/risa.13805. [24] Liu, J., Gong, E., & Wang, X. (2022). Economic benefits of construction waste recycling enterprises under tax incentive policies. Environmental Science and Pollution Research, 29(9), 12574–12588. doi:10.1007/s11356-021-13831-8. [25] Qian, H., Xu, S., Cao, J., Ren, F., Wei, W., Meng, J., & Wu, L. (2021). Air pollution reduction and climate co-benefits in China’s industries. Nature Sustainability, 4(5), 417–425. doi:10.1038/s41893-020-00669-0. [26] Phrommarat, B., & Oonkasem, P. (2021). Sustainable pineapple farm planning based on eco-efficiency and income risk: A comparison of conventional and integrated farming systems. Applied Ecology and Environmental Research, 19(4), 2701–2717. doi:10.15666/aeer/1904_27012717. [27] Guo, L., Cao, Y., Su, Q., Liu, T., & Tseng, M. L. (2023). Identifying the evolution of ecological poverty alleviation efficiency and its influencing factors: evidence from counties in Northeast China. Environmental Science and Pollution Research, 30(23), 64078–64093. doi:10.1007/s11356-023-26783-y. [28] Li, J., Chen, L., Chen, Y., & He, J. (2022). Digital economy, technological innovation, and green economic efficiency— Empirical evidence from 277 cities in China. Managerial and Decision Economics, 43(3), 616–629. doi:10.1002/mde.3406. [29] Zhao, H., Liu, G., You, S., Camargo, F. V. A., Zavelani-Rossi, M., Wang, X., Sun, C., Liu, B., Zhang, Y., Han, G., Vomiero, A., & Gong, X. (2021). Gram-scale synthesis of carbon quantum dots with a large Stokes shift for the fabrication of eco-friendly and high-efficiency luminescent solar concentrators. Energy and Environmental Science, 14(1), 396–406. doi:10.1039/d0ee02235g. [30] Oh, C. (2023). Exploring the Way to Harmonize Sustainable Development Assessment Methods in Article 6.2 Cooperative Approaches of the Paris Agreement. Green and Low-Carbon Economy, 1(3), 121–129. doi:10.47852/bonviewglce32021065. [31] Wang, Y., Liu, Y., Feng, W., & Zeng, S. (2023). Waste Haven Transfer and Poverty-Environment Trap: Evidence from EU. Green and Low-Carbon Economy, 1(1), 41–49. doi:10.47852/bonviewglce3202668. [32] Lee, C. C., & Du, L. (2024). Can green finance improve eco-efficiency? New Insights from China. Environmental Science and Pollution Research, 31(28), 40976–40994. doi:10.1007/s11356-024-33832-7. ctual value