BANGLADESH JOURNAL OF MULTIDISCIPLINARY SCIENTIFIC RESEARCH 9(2) (2024), 45-54 45 MULTIDISCIPLINARY SCIENTIFIC RESEARCH BJMSR VOL 9 NO 2 (2024) P-ISSN 2687-850X E-ISSN 2687-8518 Available online at https://www.cribfb.com Journal homepage: https://www.cribfb.com/journal/index.php/BJMSR Published by CRIBFB, USA VADER SENTIMENT ANALYSIS ON TWITTER: PREDICTING PRICE TRENDS AND DAILY RETURNS IN INDIA’S STOCK MARKET Amlan Baruah (a)1 Banajit Changkakati (b) (a)Research Scholar, Department of Business Administration, Gauhati University, Guwahati, India; E-mail: baruahamlan0@gmail.com (b)Associate Professor, Department of Business Administration, Gauhati University, Guwahati, India; E-mail: banajitc@gauhati.ac.in A R T I C L E I N F O Article History: Received: 11th May 2024 Reviewed & Revised: 11th May to 20th July 2024 Accepted: 25th July 2024 Published: 28th July 2024 Keywords: Stock Trend Prediction, Sentiment Analysis, Text Mining, Stock Market, Indian Stock Market, Machine Learning, VADER Algorithm JEL Classification Codes: G14, G17, C55, D85 Peer-Review Model: External peer review was done through double-blind method A B S T R A C T The study explores the effectiveness of sentiment analysis in predicting stock price movements, specifically focusing on the Indian Stock Market. The study investigates the reliability of social media sentiment analysis in financial markets and its implications for investors and traders. The research utilizes a sample of Twitter data comprising tweets containing hashtags related to the State Bank of India (SBI), used as a representative sample of the broader Indian Stock Market, collected from January 2021 to February 2024. The Valence Aware Dictionary for Sentiment Reasoning (VADER) algorithm was employed to analyse the sentiment of the Twitter data. Machine learning methods, including Random Forest, XGBoost, and AdaBoost, were used to integrate sentiment scores with technical indicators for predicting stock price trends. The results reveal that using only sentiment analysis achieved an accuracy of around 60% in predicting stock price direction. However, this accuracy increased to 70% with the AdaBoost method, 79% with the XGBoost method, and 82% with the Random Forest method combined with technical indicators while increasing the F1 scores from 0.4 to 0.8 in all three methods. Integrating sentiment analysis with technical indicators enhances financial market predictions by combining real- time investor sentiment with empirical historical data, leading to more accurate and adaptive trading strategies. Sentiment score was found to have a strong positive correlation with positive daily returns compared to negative daily returns, indicating that higher positive sentiment is associated with increased returns. Although negative sentiment exhibits a statistically significant correlation with daily returns, it shows a weaker positive association. © 2024 by the authors. Licensee CRIBFB, USA. This open-access article is distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0). INTRODUCTION In recent times, amidst the surge of social media interaction in the digital era, vast amounts of data covering diverse topics are consistently generated and deposited within social media platforms. This reservoir of data stands as the contemporary equivalent of a goldmine, brimming with valuable information. Within this landscape, Twitter sentiment analysis has emerged as a potent instrument in stock prediction. By harnessing the extensive real-time data flow on platforms such as X (formerly Twitter), analysts can delve into sentiments expressed in tweets related to specific stocks or financial markets, thereby uncovering valuable insights into investor sentiment, market trends, and potential price movements. This innovative approach to stock forecasting is based on the extraction, measurement, and analysis of sentiment from tweets using machine learning and natural language processing techniques. Whether positive or negative, these sentiments can serve as early indicators of shifts in market sentiment, empowering traders and investors to make more astute decisions. Integrating sentiment analysis with machine learning and natural language processing methodologies has revolutionized stock forecasting strategies. Whether identifying positive or negative sentiments, these analyses empower traders and investors to enhance decision-making processes, complementing traditional financial analysis methods. Recently, several studies have explored using VADER (Valence Aware Dictionary and Sentiment Reasoner) for sentiment analysis in stock price prediction. VADER, a lexicon and rule-based tool, has shown superior accuracy in analysing sentiments from news headlines and social media compared to traditional machine learning algorithms (Ekaputri & Akbar, 2022; Soni & Mathur, 2023). Researchers have combined VADER with other techniques like LSTM to create 1Corresponding author: ORCID ID: 0009-0002-0353-0990 © 2024 by the authors. Hosting by CRIBFB. Peer review is the responsibility of CRIBFB, USA. https://doi.org/10.46281/bjmsr.v9i2.2226 To cite this article: Baruah, A., & Changkakati, B. (2024). VADER SENTIMENT ANALYSIS ON TWITTER: PREDICTING PRICE TRENDS AND DAILY RETURNS IN INDIA’S STOCK MARKET. Bangladesh Journal of Multidisciplinary Scientific Research, 9(2), 45-54. https://doi.org/10.46281/bjmsr.v9i2.2226 http://creativecommons.org/licenses/by/4.0/) http://creativecommons.org/licenses/by/4.0/) https://www.openaccess.nl/en https://doi.org/10.46281/bjmsr.v9i2.2226 https://orcid.org/0009-0002-0353-0990 https://orcid.org/0000-0002-1802-8891 Baruah & Changkakati, Bangladesh Journal of Multidisciplinary Scientific Research 9(2) (2024), 45-54 46 hybrid models for improved stock price forecasting (Dutta et al., 2021). These approaches consider both time-series data and sentiment analysis to predict intra-day stock movements. Some studies have modified VADER by incorporating financial lexicons to enhance its performance in the financial domain (Ekaputri & Akbar, 2022). The integration of VADER- based sentiment analysis with traditional stock metrics has also been explored, resulting in accurate stock recommendation systems (Rao et al., 2022). This paper follows this general direction of integrating sentiment analysis and time-series data in the financial markets. Despite advancements in sentiment analysis for stock prediction, a critical gap persists in understanding whether sentiment analysis offers more dependable signals compared to traditional methods when applied independently or in conjunction. This gap warrants investigation to ascertain sentiment analysis's reliability and comparative advantage in forecasting price trends and daily return fluctuations within the Indian Stock Market. The primary objective of this study is to ascertain the effectiveness of sentiment analysis in predicting price trends and daily return fluctuations within the Indian Stock Market. Specifically, we aim to evaluate whether sentiment analysis provides more dependable signals compared to traditional methods when used singularly or combined with other methods. To achieve this, we utilize tweets data and financial data pertaining to the State Bank of India (SBI) as a representative sample of the broader Indian Stock Market. The remainder of the paper is organized as follows: The second section provides a comprehensive literature review of several key previous studies and presents the hypotheses. The third section describes the data and methodologies used to reach our research conclusions. The fourth section presents our experiments' results, while the fifth section offers a detailed discussion of our findings. The conclusion summarizes the research and its key points in the final section. LITERATURE REVIEW In the rapidly evolving landscape of stock market forecasting, sentiment analysis has emerged as a crucial tool for understanding and predicting market movements. The literature review explores various themes surrounding this field, emphasizing the development and evaluation of sentiment analysis tools and their application in financial contexts. Studies highlight the significant impact of sentiment analysis on stock market prediction, showcasing how sentiment extracted from social media platforms like Twitter can provide valuable insights into investor behavior and stock returns. One prevalent theme revolves around developing and evaluating sentiment analysis tools tailored for specific contexts. Hutto and Gilbert (2014) introduce (Valence Aware Dictionary and Sentiment Reasoner (VADER), a sentiment analysis tool designed for microblog-like platforms, showcasing its superiority over existing tools through a combination of qualitative and quantitative methods. The authors emphasize the importance of human expertise in the tool's development process, leading to remarkable results in sentiment analysis within computer science. Another significant theme centers on the impact of sentiment analysis on stock market prediction. Feng et al. (2017) investigate the relationship between market sentiment and the effectiveness of technical trading approaches, finding that technical indicators perform better during periods of high sentiment, particularly for small stocks. Similarly, Deng et al. (2011) propose a stock price prediction model that integrates features from both time series data and social networks, outperforming traditional methods in predicting stock prices. These studies highlight the potential of sentiment analysis to improve the accuracy of stock market forecasts. Sentiment analysis on social media significantly impacts stock trends by providing valuable insights into investor sentiments and predicting stock returns. Studies like those by Chang et al. (2021) and Gu and Kurov (2020) demonstrate that sentiment extracted from platforms like Twitter can predict stock returns without subsequent reversals, offering new information about analyst recommendations, price targets, and quarterly earnings. Additionally, research by Chen et al. (2022) shows that investor sentiment, classified by theme, is positively correlated with stock excess return, with different themes exerting varying degrees of influence on short and long-term trends. Integrating sentiment analysis with other data sources, such as news articles and historical stock data, as proposed by Ray et al. (2021) and Ho and Huang (2021), can enhance the accuracy of stock trend predictions by capturing nonlinear structures and anomalies in the market, ultimately aiding in making more informed investment decisions. Applying sentiment analysis in financial domains is another focal point of the literature. G. Wang et al. (2014) assess the impact of content on social investment platforms, demonstrating the outperformance of Seeking Alpha articles in stock market returns over a baseline market. Njølstad et al. (2014) address limitations in sentiment analysis within news articles, proposing feature categories and machine learning methods to enhance classification precision. Additionally, Bhardwaj et al. (2015) explore the influence of internet-based technologies on the Indian stock market, underscoring the significance of sentiment analysis in forecasting stock prices. Utilizing social media data for stock market analysis emerges as a prominent subtheme. Nguyen et al. (2015) have demonstrated that incorporating topic-specific sentiments from social media can enhance stock price movement prediction models, outperforming historical price-based methods. Similarly, Batra and Daudpota (2018) apply sentiment analysis to predict stock movements using tweets related to Apple products, revealing a positive correlation between sentiments expressed in tweets and market data. Various methods are employed across the studies to conduct sentiment analysis and predict stock market movements. These methods include machine learning algorithms like Support Vector Machines (SVM), Decision Trees, Random Forests, and Multiple Kernel Learning regression frameworks. Additionally, qualitative analyses, empirical analyses, and correlation studies are conducted to evaluate the effectiveness of sentiment analysis tools and predictive models. Baruah & Changkakati, Bangladesh Journal of Multidisciplinary Scientific Research 9(2) (2024), 45-54 47 Machine learning algorithms predict stock market trends using social media trends by analyzing alternative data sources like social media commentary, news articles, and sentiment analytics (Ashtiani & Raahemi, 2023; Dong et al., 2021; Gülmez, 2023; Hansen & Borch, 2022; Sharaf et al., 2023). These algorithms utilize sentiment analysis of COVID-19 news, text mining on social media data, and GPS data to extract insights for stock price prediction (Ashtiani & Raahemi, 2023; Sharaf et al., 2023). Techniques such as LSTM models optimized by metaheuristic algorithms like ARO, dynamic predictor selection algorithms, and text mining are employed to process social media data and predict stock movements accurately (Ashtiani & Raahemi, 2023; Dong et al., 2021; Gülmez, 2023). By harnessing the power of machine learning and text mining on social media data, these algorithms can provide valuable insights for investors and traders in predicting stock market trends based on social media trends and sentiment analysis. The use of these diverse methods underscores the interdisciplinary nature of sentiment analysis in finance and highlights the need for robust analytical approaches to extract insights from large datasets. Despite the potential of sentiment analysis, contradictions in using social media models to predict the stock market lie in the challenges of user motivations, prediction quality, and the impact of external factors. While social media platforms like StockTwits can be predictive of stock performance (Bouadjenek et al., 2023), incorporating sentiments and technical indicators can enhance prediction accuracy (Z. Wang et al., 2023). However, the presence of misleading information from users with potentially ulterior motives can hinder the reliability of predictions (Bouadjenek et al., 2023). Additionally, the dynamic and nonlinear nature of stock trends requires careful consideration of various influence factors at different phases, emphasizing the need for a hybrid model that integrates social media sentiments and technical indicators for more accurate predictions (Z. Wang et al., 2023). Despite the potential of deep learning techniques and word embedding methods to forecast stock movements using social media data, the high volatility of stock markets and the randomness of events pose challenges in accurately quantifying the different influences of news and social media on stock prices (Khan et al., 2022; Kilimci & Duvar, 2020; J. Liu et al., 2020). Nevertheless, social media models are justified for predicting financial asset prices due to their ability to provide real-time data, sentiment analysis, and market efficiency insights. Studies like Polyzos et al. (2024) propose using social media as a proxy for financial information, showcasing successful forecasts for over 8000 cryptocurrencies. Additionally, research by Wang introduces a multimodal deep learning model that incorporates Twitter content to predict extreme price fluctuations in Bitcoin, demonstrating the impact of social media on asset prices (Zou & Herremans, 2023). Furthermore, De Arriba-Perez et al. (2020) highlight the value of detecting positive predictions in tweets to support investors' decision- making, emphasizing the importance of sentiment analysis from micro-blogging sources like Twitter. Researchers can enhance stock movement predictions and improve investment decision-making processes by leveraging social media data, including opinions, sentiments, and news shared online (Mehta et al., 2021; Sawhney et al., 2020). The primary unresolved issues in sentiment analysis models that impede their accurate prediction of stock prices include the lack of transparency and explainability in deep neural network (DNN)-based methods (Mei et al., 2023), the challenge of balancing performance and resource consumption in multimodal sentiment analysis, especially when facing missing modalities, and the difficulty in identifying erroneous predictions in sentiment analysis models prior to deployment (Z. Liu et al., 2021). Incorporating financial news data alongside stock fundamental features has enhanced prediction accuracy, indicating the importance of considering textual data in stock price forecasting (Dahal et al., 2023; Rubi et al., 2022). By addressing these issues through enhanced transparency, resource optimization, and error detection mechanisms, sentiment analysis models can improve their ability to accurately predict stock prices in the highly volatile and complex stock market environment. Despite these challenges, the justification for continuing research in this area is robust. Social media platforms provide real-time data and market efficiency insights invaluable for financial forecasting. Studies such as Polyzos et al. (2024) & Zou and Herremans (2023) demonstrate the successful use of social media in predicting asset prices, underscoring the importance of this data source. Moreover, incorporating financial news and textual data alongside traditional stock features has enhanced prediction accuracy (Dahal et al., 2023). Given these considerations, this study aims to explore the use of advanced machine learning algorithms— Random Forest, Extreme Gradient Boosting (XGBoost), and Adaptive Boosting (AdaBoost)—in conjunction with the VADER sentiment analysis tool to predict stock market prices. These algorithms, known for their robustness and ability to handle complex datasets, will leverage the nuanced sentiment data extracted by VADER from social media platforms like X (formerly Twitter). By addressing the identified challenges and building on the strengths of previous research, this study seeks to advance the field of sentiment-driven financial forecasting, offering more accurate and interpretable predictions for stock market movements. The primary objective of this study is to evaluate the effectiveness of sentiment analysis in predicting stock price movements within the Indian stock market, both independently and when combined with technical indicators, taking State Bank of India (SBI) stock as the representative sample for this analysis. Additionally, a secondary objective is to assess the effectiveness of sentiment analysis in forecasting daily return fluctuations. The following research hypothesis is developed to be tested during the research- H01: Sentiment analysis does not provide more dependable signals when combined with technical analysis when predicting price trends of the State Bank of India (SBI) stock. H02: Sentiment analysis does not effectively predict daily return fluctuations of the State Bank of India (SBI) stock. Baruah & Changkakati, Bangladesh Journal of Multidisciplinary Scientific Research 9(2) (2024), 45-54 48 MATERIALS AND METHODS Material Selection The dataset in this study comprises Twitter hashtag data gathered through the utilization of the X (formerly Twitter) API and data provided by the third-party Twitter data vendor 'TweetBinder'. Specifically, the dataset encompasses tweets featuring the hashtags #SBIN and #SBI, amounting to a total of 3087 tweets. These tweets were amassed over a comprehensive period spanning three years, commencing in January 2021 and concluding in February 2024. Notably, the dataset encapsulates a rich array of insights derived from Twitter users' discussions pertaining to the State Bank of India (SBI). Data from 654 unique days was incorporated throughout the analysis, ensuring a robust and diverse representation of temporal dynamics within the dataset. The financial data of the stock 'SBI' is derived from the use of the 'YAHOO_FIN' library in Python, which derives data from the website- 'www.yahoofinance.com'. The technical analysis used two technical indicators, the 'Relative Strength Index'(RSI) and the 'On-Balance Volume'(OBV), to predict the price signal. The daily price signals are derived from the ‘ adjusted close’ column. The price signals are given by: Price Signal = Where pt is the current price and pt-1 is the previous price. Preprocessing Prior to analysis, the dataset underwent meticulous scrutiny to filter out spam tweets. Spam tweets were identified as promoting stock purchases with promises of immediate gains, needing more substantive market sentiment. After manual review, 3087 tweets across 654 trading days were deemed suitable for further analysis, ensuring the integrity and reliability of the dataset. Measures and Covariates The study utilized primary outcome measures such as sentiment scores derived from the VADER algorithm applied to social media data (tweets). Sentiment scores were categorized into positive, negative, and neutral sentiments based on numerical values ranging from -1 (most negative) to 1 (most positive), with 0 representing neutral sentiment. Secondary outcome measures included financial metrics like the daily returns of the SBI stock, calculated as the percentage change in stock price from one trading day to the next. Covariates considered in the analysis included technical indicators commonly used in financial analysis, such as the Relative Strength Index (RSI) and On-Balance Volume (OBV). These covariates were integrated into the analysis to assess their combined predictive power with sentiment scores on stock price movements. Research Design The research design employed in this study was a correlational research design. It focused on analyzing sentiment trends in social media data and their correlation with stock market performance without manipulating variables or creating experimental conditions. Sentiment analysis was conducted using the VADER algorithm to categorize tweets as positive, negative, or neutral, reflecting public sentiment towards the SBI stock. Sampling Procedures Data collection utilized secondary data obtained through judgmental sampling. This sampling approach involved selecting tweets related to SBI stock from the dataset based on their relevance and authenticity, ensuring the inclusion of diverse perspectives and market conditions. VADER Sentiment Score The Valence Aware Dictionary for Sentiment Reasoning (VADER) (Hutto & Gilbert, 2014) stands as a significant advancement in Natural Language Processing (NLP), designed to gauge the polarity and intensity of sentiment expressions. Leveraging a comprehensive lexicon comprising words, phrases, emoticons, and acronyms, each meticulously rated by human annotators for polarity and intensity, Vader employs a sophisticated algorithmic framework augmented by grammatical rules. These rules effectively handle linguistic nuances such as negations and intensifiers, ensuring a nuanced sentiment assessment across various textual contexts. Vader, integrated into the widely-used Natural Language Toolkit (NLTK) within the Python programming language, boasts an expansive lexicon encompassing approximately 7,500 sentiment features, with unlisted words defaulted to a neutral sentiment classification. In this study, sentiment analysis of selected stocks was conducted utilizing the Natural Language Toolkit (NLTK) within the Python programming language. Each tweet was subject to individual analysis by the Python program package, generating two distinct sentiment scores: a positive sentiment score and a negative sentiment score. Subsequently, the final sentiment score for each tweet was derived through a weighted average calculation, wherein the contribution of each word within the tweet's textual composition was considered. In the case of multiple tweets in a single day, the average 'Sentiment Score' of the tweets is taken as the 'Sentiment Score' of the day. { 1 𝑖𝑓 𝑝𝑡 > 𝑝𝑡−1 0 𝑖𝑓 𝑝𝑡 ≤ 𝑝𝑡−1 Baruah & Changkakati, Bangladesh Journal of Multidisciplinary Scientific Research 9(2) (2024), 45-54 49 Sample Validation The dataset was divided into training and testing sets in an 80:20 ratio using randomization to validate the study's findings. Additionally, numerical values used in both the training and testing phases were normalized using min-max normalization. This normalization technique standardizes the scale of data values, reducing potential biases arising from differing data ranges and enhancing the reliability of analytical results. RESULTS The following results were obtained after using relevant libraries like ‘scikit-learn’, ‘vaderSentiment’, and NLTK in Python, executed through the Google Colab environment running on cloud-based GPUs provided by Google. VADER Sentiment Scores The tweets are classified as positive, negative, or neutral based on the sentiment score, which ranges from 1 to -1, with 1 being the most positive and -1 being the most negative. A score of 0 is taken as neutral sentiment. In our analysis, the distribution of tweets is represented by the pie diagram given below: Figure 1. The Distribution of Positive, Negative, and Neutral Tweets in the dataset Reliability Analysis The reliability analysis was conducted by scrutinizing price trend data, represented as binary values (1 for upward trend, 0 for otherwise), in conjunction with sentiment scores derived from the VADER algorithm, ranging from -1 to 1. Initially, the price trend was evaluated solely based on sentiment scores with supervised machine learning algorithms. Subsequently, the price trend underwent analysis incorporating machine learning techniques applied to technical indicators, followed by the inclusion of sentiment scores corresponding to each respective day. Comparative assessments were made across the three scenarios, examining accuracy scores and F1 scores to discern their relative efficacy. Our experiment uses three supervised machine learning methods: Random Forest Classifier, XGBoost, and AdaBoost. Table 1. Accuracy and F1 Scores of Sentiment Analysis using different methods and combinations Method ML Method Accuracy F1 score Only Sentiment Scores Random Forest 0.601 0.452 XGBoost 0.618 0.452 AdaBoost 0.511 0.452 Only Technical Indicators Random Forest 0.756 0.738 XGBoost 0.802 0.794 AdaBoost 0.664 0.793 Sentiment Scores + Technical Indicators Random Forest 0.824 0.806 XGBoost 0.794 0.807 AdaBoost 0.702 0.807 *Accuracy and F1 relate to machine learning algorithms. For more details, see Note A and Note B. The table presents a comparative analysis of sentiment analysis methods utilizing sentiment scores alone, technical indicators alone, and a combination of both. Across various machine learning algorithms such as Random Forest, XGBoost, and AdaBoost, the incorporation of sentiment scores alongside technical indicators consistently demonstrates an improvement in accuracy. Notably, adding sentiment scores leads to significant enhancements in accuracy compared to using technical indicators alone, with the combined approach yielding the highest accuracy scores. Despite this boost in accuracy, the F1 scores, which reflect the balance between precision and recall, remain consistent or slightly increase when sentiment scores are integrated, indicating that the combination maintains the model's ability to correctly classify sentiment without sacrificing performance metrics. Baruah & Changkakati, Bangladesh Journal of Multidisciplinary Scientific Research 9(2) (2024), 45-54 50 Since our first null hypothesis (H01) was 'Sentiment analysis does not provide more dependable signals when combined with technical analysis when predicting price trends of the State Bank of India (SBI) stock', we reject the null hypothesis and accept the alternative hypothesis that sentiment analysis, when combined with technical indicators, increases the predictive ability to know the price trend. The results are validated through the use of separate train and test data. These findings underscore the significance of incorporating sentiment scores alongside traditional technical indicators, offering a more robust and accurate approach to sentiment analysis that can potentially enhance decision-making processes in financial markets. Effect of Sentiments on Daily Returns To determine the effect of sentiments on daily returns, we calculate the correlation of the stock's daily returns with the overall sentiment score. Additionally, the analysis is conducted separately for the days when the daily returns are either positive or negative. Table 2. Correlation between Sentiment Score and Daily Returns The table shows the correlation between sentiment scores and daily returns of SBI in the National Stock Exchange (NSE) of India, delineated across distinct sentiment directions: positive, negative, and any direction. The correlation test results enable us to reject the null hypothesis (H02) that ‘Sentiment analysis does not effectively predict daily return fluctuations of the State Bank of India (SBI) stock’. Since the p-value is less than 0.05, indicating a significant correlation between Twitter sentiments and daily return fluctuations in any three conditions. Therefore, we reject the null hypothesis and conclude that there is a significant correlation between sentiment analysis and daily returns. DISCUSSIONS The classification of tweets into positive, negative, and neutral categories based on the VADER sentiment scores provides an insightful overview of the sentiment landscape within the dataset. By categorizing the sentiment scores into these three distinct groups, we better understand the prevailing emotional tones expressed in the tweets. This foundational step is critical as it sets the stage for deeper analyses, allowing us to track sentiment trends over time and correlate them with market movements. Visualizing this distribution through a pie chart further aids in illustrating the proportion of each sentiment type, highlighting the dominance or scarcity of particular sentiments within the dataset. The reliability analysis underscores the importance of integrating sentiment scores with traditional technical indicators in enhancing the predictive accuracy of financial models. The comparative evaluation using Random Forest, XGBoost, and AdaBoost algorithms reveals that models incorporating both sentiment scores and technical indicators consistently outperform those relying on a single data source. This finding is significant as it demonstrates that the addition of sentiment scores can provide a richer, more nuanced understanding of market dynamics, thereby improving model accuracy. The accuracy of the 3 machine learning methods received a boost of an average of 19.6 % when used in conjunction with technical indicators. The slight improvements in F1 scores suggest that the enhanced accuracy does not come at the expense of the model's precision or recall capabilities. This holistic approach to financial modeling, which leverages both sentiment analysis and technical indicators, offers a robust framework for market prediction, potentially leading to better-informed investment decisions. Analyzing the correlation between sentiment scores and daily returns provides compelling evidence of the impact of market sentiment on stock performance. The strong positive correlation between positive sentiment and daily returns indicates that higher positive sentiment is associated with increased returns, a statistically significant finding. This relationship highlights the influence of public sentiment on investor behavior and market outcomes. On the other hand, negative sentiment, while also showing a significant correlation with daily returns, exhibits a weaker positive association. This suggests that while negative sentiments do affect market performance, their impact is less pronounced compared to positive sentiments. The modest yet significant correlation for sentiment in any direction further emphasizes the importance of sentiment analysis in understanding market dynamics. These findings align with existing literature on the subject, reinforcing the notion that market sentiment plays a crucial role in financial markets. In prior research, Pagolu et al. (2016) analyzed around 250,000 Microsoft-related tweets using N-Gram and Word2vec methods, reporting accuracies of 57% to 70%. Despite a smaller sample, our study's lowest accuracy was 62%, rising to 82% when combined with technical indicators using Random Forest. In another study, Mardjo and Choksuchat (2022) used the HyVADRF (Hybrid Valence Aware Dictionary and Sentiment Reasoner–Random Forest) and Gray Wolf Optimizer (GWO) model. VADER calculated polarity scores and classified sentiments, overcoming manual labeling weaknesses, while Random Forest acted as the supervised classifier. The researchers collected tweets, analyzed dataset sizes, and used GWO for parameter tuning, achieving a 75.29% accuracy. This study solely used sentiment analysis, while our research integrates sentiment analysis and technical analysis, which increased the accuracy score by the Random Forest method to 82.4%. When integrated with technical indicators, Sentiment analysis enhances the accuracy and reliability of financial market predictions through a synergistic approach that leverages qualitative and quantitative data. While sentiment analysis gauges market sentiment and investor emotions from textual sources such as social media, news articles, and financial Daily Return Correlation p-value Positive 0.5890248051173522 3.4133153188880307e-10 Negative 0.19155963401027593 0.036089508382569815 Any direction 0.11926219140966826 0.002250469657582959 Baruah & Changkakati, Bangladesh Journal of Multidisciplinary Scientific Research 9(2) (2024), 45-54 51 reports, technical indicators provide empirical insights derived from historical price and volume data. By combining these methodologies, analysts can capture nuanced market behaviors that neither method alone can fully discern. Sentiment analysis enriches technical analysis by offering real-time insights into investor perceptions and market psychology, which can influence trading decisions and price trends. Moreover, the incorporation of sentiment data into technical models allows for more adaptive and responsive trading strategies that reflect current market sentiment alongside historical trends, thereby improving overall forecasting accuracy and decision-making processes in financial markets. Sentiment analysis shows a stronger correlation with positive daily returns in the stock market due to investor behaviors influenced by optimism and confidence. Positive sentiment tends to amplify market momentum, leading to continued price appreciation as more investors join the bullish sentiment bandwagon. In contrast, negative sentiment, associated with fear and uncertainty, may not always trigger immediate or significant price declines, as investors may react less decisively to pessimistic signals. Investors may react more swiftly and decisively to positive news or sentiment due to the allure of potential gains. In contrast, negative sentiment may sometimes be discounted or countered by other market factors, such as long-term fundamentals or market corrections. Future research prospects include incorporating computationally intensive deep learning models with VADER for improved accuracy and contextual understanding of social media sentiments influencing market trends. Expanding data sources to include diverse platforms like forums and blogs can provide a richer dataset that better reflects market sentiment, aiding in stock price prediction. Additionally, refining VADER's sentiment classification and incorporating temporal and cross-market analysis could offer more granular insights and reveal patterns of sentiment contagion across different markets, ultimately enhancing investment decision-making. CONCLUSIONS This study utilized the VADER sentiment analysis algorithm implemented in Python's NLTK to evaluate its effectiveness in predicting stock price trends, focusing specifically on the SBI stock from January 2021 to February 2024. The analysis revealed that sentiment analysis achieved an accuracy rate of approximately 60% in predicting stock price directions independently, with substantial improvements noted when integrated with technical indicators. Among the machine learning models tested, the random forest classifier consistently outperformed XGBoost and AdaBoost, highlighting its efficacy in combining sentiment analysis with technical signals for enhanced predictive accuracy. The study's findings underscore the significant influence of sentiment on market dynamics, remarkably amplifying bullish sentiment during periods of rising prices. Conversely, bearish sentiment showed less impact during market downturns compared to bullish sentiment during upturns. Despite these insights, the study acknowledges challenges such as linguistic subjectivity, cultural nuances, and data scarcity, which limit the reliability of sentiment analysis. The contributions of this research lie in its demonstration of sentiment analysis as a valuable supplementary tool for market participants, offering enhanced insights when integrated with traditional analytical approaches. Investors, traders, and financial analysts can leverage these insights to refine decision-making strategies, combining sentiment analysis with technical analysis to better anticipate market trends and optimize investment outcomes. Theoretical implications highlight the evolving landscape of financial analysis, where sentiment analysis offers a nuanced perspective on market sentiment dynamics. Managerially, this study encourages practitioners to adopt integrated analytical frameworks incorporating sentiment analysis, enriching their understanding of market behaviour and improving decision-making in volatile market conditions. Future research could explore the robustness of sentiment analysis methodologies during periods of market volatility, examine the influence of market sentiment in intraday trading contexts, and investigate the impact of financial news on market opening conditions. Author Contributions: Conceptualization, A.B. and B.C.; Methodology, A.B.; Software, A.B.; Validation, A.B. and B.C.; Formal Analysis, A.B. and B.C.; Investigation, A.B. and B.C.; Resources, A.B. and B.C.; Data Curation, A.B., Writing – Original Draft Preparation, A.B. and B.C.; Writing – Review & Editing, A.B. and B.C.; Visualization, A.B.; Supervision, B.C.; Project Administration, A.B. and B.C.; Funding Acquisition, A.B. and B.C. The author has read and agreed to the published version of the manuscript. Institutional Review Board Statement: Ethical review and approval were waived for this study because the research does not involve vulnerable groups or sensitive issues. Funding: The authors received no direct funding for this research. Acknowledgments: Not Applicable. Informed Consent Statement: Informed consent was obtained from all subjects involved in the study. Data Availability Statement: The data presented in this study are available upon request from the corresponding author. Due to restrictions, they are not publicly available. Conflicts of Interest: The authors declare no conflict of interest. NOTES Note A: Accuracy Score Accuracy is the ratio of correctly predicted instances to the total instances in the dataset. Accuracy = Note B: F1 Score The F1 score is the harmonic mean of precision and recall. It is a single metric that combines both precision and recall. F1 Score = Number of Correct Predictions Total Number of Predictions 2 ∗ Precision × Recall Precision + Recal Baruah & Changkakati, Bangladesh Journal of Multidisciplinary Scientific Research 9(2) (2024), 45-54 52 Where: Precision is the ratio of correctly predicted positive observations to the total predicted positive observations. Precision = Recall (also known as sensitivity or true positive rate) is the ratio of correctly predicted positive observations to all the actual positive. Recall = APPENDICES Appendix A: Calculation of VADER Scores Calculation of VADER scores of a tweet may be summarized with the help of the following example. Example tweet – "Kudos to SBI for making banking more accessible and inclusive with the revamped YONO app. Now, everyone can enjoy the benefits of digital banking! 🌐 #AccessibleBanking #InclusiveServices" Step-by-Step Calculation: 1. Tokenize the Tweet: "Kudos", "to", "SBI", "for", "making", "banking", "more", "accessible", "and", "inclusive", "with", "the", "revamped", "YONO", "app.", "Now,", "everyone", "can", "enjoy", "the", "benefits", "of", "digital", "banking!", "🌐", "#AccessibleBanking", "#InclusiveServices" 2. Identify Sentiment Words and Scores: According to the VADER lexicon, the words and their scores are: "Kudos": +2.0, "accessible": +1.5, "inclusive": +1.5, "enjoy": +2.0, "benefits": +1.5 3. Adjust for Modifiers and Punctuation: There are no specific intensifiers or negations in this tweet, but there is an exclamation mark. Exclamation marks ("!") can amplify the sentiment score. One exclamation mark typically adds 0.292 to the sentiment intensity. Let's adjust the scores: o "Kudos": +2.0 (no modifier) o "accessible": +1.5 (no modifier) o "inclusive": +1.5 (no modifier) o "enjoy": +2.0 (no modifier) o "benefits": +1.5 (no modifier) o Adding the exclamation mark: Let's assume it adds 0.292 to the total positive sentiment score. 4. Sum Up the Scores: o Positive sentiment: 2.0 (Kudos) + 1.5 (accessible) + 1.5 (inclusive) + 2.0 (enjoy) + 1.5 (benefits) + 0.292 (exclamation mark) = 8.792 o Negative sentiment: 0 (no negative words) o Neutral sentiment: Non-sentiment words are counted as neutral. For simplicity, let's count all other words as neutral. 5. Normalize the Scores: o The positive, neutral, and negative scores are normalized by the total number of words. Final Scores Calculation:  Positive Score: Sum of positive sentiment scores divided by the number of sentiment words: o Positive Score = 8.7925/5 = 1.7584 (considering five sentiment-bearing words: "Kudos", "accessible", "inclusive", "enjoy", and "benefits")  Negative Score: There are no negative sentiment words, so: o Negative Score = 0  Neutral Score: If there are 21 neutral words: o Neutral Score = 21/26 ≈ 0.8077  Compound Score: The compound score is a normalized score ranging from -1 to 1, calculated using a formula that incorporates the positive and negative scores with a normalization factor. This score is not as straightforward to calculate manually because VADER uses a specific formula involving a sigmoid function to compute it. In our example:  Positive: 1.7584  Negative: 0 True Positive True Positives + False Positives True Positives True Positives + False Negatives Baruah & Changkakati, Bangladesh Journal of Multidisciplinary Scientific Research 9(2) (2024), 45-54 53  Neutral: 0.8077 (normalized by total words)  Compound: Using VADER's internal algorithm, which typically involves summing the normalized positive and negative scores and applying a normalization factor. These calculations are done manually to show how VADER evaluates the sentiment scores by considering the sentiment intensity of individual words, applying adjustments for modifiers, and normalizing the results. To get the exact compound score, it would be best to use the VADER library directly, as it incorporates more nuanced adjustments and normalization. REFERENCES Ashtiani, M. N., & Raahemi, B. (2023). News-based intelligent prediction of financial markets using text mining and machine learning: A systematic literature review. Expert Systems with Applications, 217, 119509. https://doi.org/10.1016/j.eswa.2023.119509 Batra, R., & Daudpota, S. M. (2018). Integrating StockTwits with sentiment analysis for better prediction of stock price movement. 2018 International Conference on Computing, Mathematics and Engineering Technologies (iCoMET), 1–5. https://doi.org/10.1109/ICOMET.2018.8346382 Bhardwaj, A., Narayan, Y., Vanraj, Pawan, & Dutta, M. (2015). Sentiment Analysis for Indian Stock Market Prediction Using Sensex and Nifty. Procedia Computer Science, 70, 85–91. https://doi.org/10.1016/j.procs.2015.10.043 Bouadjenek, M. R., Sanner, S., & Wu, G. (2023). A User-Centric Analysis of Social Media for Stock Market Prediction. ACM Transactions on the Web, 17(2), 1–22. https://doi.org/10.1145/3532856 Chang, J., Tu, W., Yu, C., & Qin, C. (2021). Assessing dynamic qualities of investor sentiments for stock recommendation. Information Processing & Management, 58(2), 102452. https://doi.org/10.1016/j.ipm.2020.102452 Chen, M., Guo, Z., Abbass, K., & Huang, W. (2022). Analysis of the impact of investor sentiment on stock price using the latent dirichlet allocation topic model. Frontiers in Environmental Science, 10, 1068398. https://doi.org/10.3389/fenvs.2022.1068398 Dahal, K. R., Pokhrel, N. R., Gaire, S., Mahatara, S., Joshi, R. P., Gupta, A., Banjade, H. R., & Joshi, J. (2023). A comparative study on effect of news sentiment on stock price prediction with deep learning architecture. PLOS ONE, 18(4), e0284695. https://doi.org/10.1371/journal.pone.0284695 De Arriba-Perez, F., Garcia-Mendez, S., Regueiro-Janeiro, J. A., & Gonzalez-Castano, F. J. (2020). Detection of Financial Opportunities in Micro-Blogging Data With a Stacked Classification System. IEEE Access, 8, 215679–215690. https://doi.org/10.1109/ACCESS.2020.3041084 Deng, S., Mitsubuchi, T., Shioda, K., Shimada, T., & Sakurai, A. (2011). Combining Technical Analysis with Sentiment Analysis for Stock Price Prediction. 2011 IEEE Ninth International Conference on Dependable, Autonomic and Secure Computing, 800–807. https://doi.org/10.1109/DASC.2011.138 Dong, S., Wang, J., Luo, H., Wang, H., & Wu, F.-X. (2021). A dynamic predictor selection algorithm for predicting stock market movement. Expert Systems with Applications, 186, 115836. https://doi.org/10.1016/j.eswa.2021.115836 Dutta, A., Pooja, G., Jain, N., Panda, R. R., & Nagwani, N. K. (2021). A Hybrid Deep Learning Approach for Stock Price Prediction. In A. Joshi, M. Khosravy, & N. Gupta (Eds.), Machine Learning for Predictive Analysis (Vol. 141, pp. 1–10). Springer Singapore. https://doi.org/10.1007/978-981-15-7106-0_1 Ekaputri, A. P., & Akbar, S. (2022). Financial News Sentiment Analysis using Modified VADER for Stock Price Prediction. 2022 9th International Conference on Advanced Informatics: Concepts, Theory and Applications (ICAICTA), 1–6. https://doi.org/10.1109/ICAICTA56449.2022.9932925 Feng, S., Wang, N., & Zychowicz, E. J. (2017). Sentiment and the Performance of Technical Indicators. The Journal of Portfolio Management, 43(3), 112–125. https://doi.org/10.3905/jpm.2017.43.3.112 Gu, C., & Kurov, A. (2020). Informational role of social media: Evidence from Twitter sentiment. Journal of Banking & Finance, 121, 105969. https://doi.org/10.1016/j.jbankfin.2020.105969 Gülmez, B. (2023). Stock price prediction with optimized deep LSTM network with artificial rabbits optimization algorithm. Expert Systems with Applications, 227, 120346. https://doi.org/10.1016/j.eswa.2023.120346 Hansen, K. B., & Borch, C. (2022). Alternative data and sentiment analysis: Prospecting non-standard data in machine learning-driven finance. Big Data & Society, 9(1), 205395172110707. https://doi.org/10.1177/20539517211070701 Ho, T.-T., & Huang, Y. (2021). Stock Price Movement Prediction Using Sentiment Analysis and CandleStick Chart Representation. Sensors, 21(23), 7957. https://doi.org/10.3390/s21237957 Hutto, C., & Gilbert, E. (2014). Vader: A parsimonious rule-based model for sentiment analysis of social media text. Proceedings of the International AAAI Conference on Web and Social Media, 8(1), 216–225. https://doi.org/10.1609/icwsm.v8i1.14550 Khan, W., Ghazanfar, M. A., Azam, M. A., Karami, A., Alyoubi, K. H., & Alfakeeh, A. S. (2022). Stock market prediction using machine learning classifiers and social media news. Journal of Ambient Intelligence and Humanized Computing, 13(7), 3433–3456. https://doi.org/10.1007/s12652-020-01839-w Kilimci, Z. H., & Duvar, R. (2020). An Efficient Word Embedding and Deep Learning Based Model to Forecast the Direction of Stock Exchange Market Using Twitter and Financial News Sites: A Case of Istanbul Stock Exchange (BIST 100). IEEE Access, 8, 188186–188198. https://doi.org/10.1109/ACCESS.2020.3029860 Baruah & Changkakati, Bangladesh Journal of Multidisciplinary Scientific Research 9(2) (2024), 45-54 54 Liu, J., Lin, H., Yang, L., Xu, B., & Wen, D. (2020). Multi-Element Hierarchical Attention Capsule Network for Stock Prediction. IEEE Access, 8, 143114–143123. https://doi.org/10.1109/ACCESS.2020.3014506 Liu, Z., Guo, Y., & Mahmud, J. (2021). When and Why a Model Fails? A Human-in-the-loop Error Detection Framework for Sentiment Analysis. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Industry Papers, 170–177. https://doi.org/10.18653/v1/2021.naacl-industry.22 Mardjo, A., & Choksuchat, C. (2022). HyVADRF: Hybrid VADER–Random Forest and GWO for Bitcoin Tweet Sentiment Analysis. IEEE Access, 10, 101889–101897. https://doi.org/10.1109/ACCESS.2022.3209662 Mehta, P., Pandya, S., & Kotecha, K. (2021). Harvesting social media sentiment analysis to enhance stock market prediction using deep learning. PeerJ Computer Science, 7, e476. https://doi.org/10.7717/peerj-cs.476 Mei, X., Zhou, Y., Zhu, C., Wu, M., Li, M., & Pan, S. (2023). A disentangled linguistic graph model for explainable aspect- based sentiment analysis. Knowledge-Based Systems, 260, 110150. https://doi.org/10.1016/j.knosys.2022.110150 Nguyen, T. H., Shirai, K., & Velcin, J. (2015). Sentiment analysis on social media for stock movement prediction. Expert Systems with Applications, 42(24), 9603–9611. https://doi.org/10.1016/j.eswa.2015.07.052 Njølstad, P. C. S., Høysæter, L. S., Wei, W., & Gulla, J. A. (2014). Evaluating Feature Sets and Classifiers for Sentiment Analysis of Financial News. 2014 IEEE/WIC/ACM International Joint Conferences on Web Intelligence (WI) and Intelligent Agent Technologies (IAT), 2, 71–78. https://doi.org/10.1109/WI-IAT.2014.82 Pagolu, V. S., Reddy, K. N., Panda, G., & Majhi, B. (2016). Sentiment analysis of Twitter data for predicting stock market movements. 2016 International Conference on Signal Processing, Communication, Power and Embedded System (SCOPES), 1345–1350. https://doi.org/10.48550/arXiv.1610.09225 Polyzos, E., Rubbaniy, G., & Mazur, M. (2024). Efficient Market Hypothesis on the blockchain: A social‐media‐based index for cryptocurrency efficiency. Financial Review, 59(3), 807–829. https://doi.org/10.1111/fire.12387 Rao, J., Ramaraju, V., Smith, J., & Bansal, A. (2022). A Sentiment Analysis Based Stock Recommendation System. 2022 IEEE Fifth International Conference on Artificial Intelligence and Knowledge Engineering (AIKE), 82–89. https://doi.org/10.1109/AIKE55402.2022.00020 Rubi, M. A., Chowdhury, S., Rahman, A. A. A., Meero, A., Zayed, N. M., & Islam, K. M. A. (2022). Fitting multi-layer feed forward neural network and autoregressive integrated moving average for Dhaka Stock Exchange price predicting. Emerging Science Journal, 6(5), 1046-1061. https://doi.org/10.28991/ESJ-2022-06-05-09 Ray, P., Ganguli, B., & Chakrabarti, A. (2021). A Hybrid Approach of Bayesian Structural Time Series With LSTM to Identify the Influence of News Sentiment on Short-Term Forecasting of Stock Price. IEEE Transactions on Computational Social Systems, 8(5), 1153–1162. https://doi.org/10.1109/TCSS.2021.3073964 Sawhney, R., Agarwal, S., Wadhwa, A., & Shah, R. R. (2020). Deep Attentive Learning for Stock Movement Prediction From Social Media Text and Company Correlations. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 8415–8426. https://doi.org/10.18653/v1/2020.emnlp-main.676 Sharaf, M., Hemdan, E. E.-D., El-Sayed, A., & El-Bahnasawy, N. A. (2023). An efficient hybrid stock trend prediction system during COVID-19 pandemic based on stacked-LSTM and news sentiment analysis. Multimedia Tools and Applications, 82(16), 23945–23977. https://doi.org/10.1007/s11042-022-14216-w Soni, J., & Mathur, K. (2023). Sentiment Analysis of News Headlines for Stock Market Prediction using VADER. 2023 3rd International Conference on Innovative Mechanisms for Industry Applications (ICIMIA), 1215–1222. https://doi.org/10.1109/ICIMIA60377.2023.10426095 Wang, G., Wang, T., Wang, B., Sambasivan, D., Zhang, Z., Zheng, H., & Zhao, B. Y. (2014). Crowds on wall street: Extracting value from social investing platforms. arXiv Preprint arXiv:1406.1137. https://doi.org/10.48550/arXiv.1406.1137 Wang, Z., Hu, Z., Li, F., Ho, S.-B., & Cambria, E. (2023). Learning-Based Stock Trending Prediction by Incorporating Technical Indicators and Social Media Sentiment. Cognitive Computation, 15(3), 1092–1102. https://doi.org/10.1007/s12559-023-10125-8 Zou, Y., & Herremans, D. (2023). PreBit—A multimodal model with Twitter FinBERT embeddings for extreme price movement prediction of Bitcoin. Expert Systems with Applications, 233, 120838. https://doi.org/10.1016/j.eswa.2023.120838 Publisher’s Note: CRIBFB stays neutral with regard to jurisdictional claims in published maps and institutional affiliations. © 2024 by the authors. Licensee CRIBFB, USA. This open-access article is distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0). Bangladesh Journal of Multidisciplinary Scientific Research (P-ISSN 2687-850X E-ISSN 2687-8518) by CRIBFB is licensed under a Creative Commons Attribution 4.0 International License. http://creativecommons.org/licenses/by/4.0). http://creativecommons.org/licenses/by/4.0/ http://creativecommons.org/licenses/by/4.0/ http://creativecommons.org/licenses/by/4.0/