Academic Journal of Science and Technology ISSN: 2771-3032 | Vol. 10, No. 1, 2024 284 Image Captioning in News Report Scenario Tianrui Liu1, *, Qi Cai2, Changxin Xu3, Bo Hong4, Jize Xiong5, Yuxin Qiao6, Tsungwei Yang7 1 Department of Electrical and Computer Engineering, University of California San Diego, USA 2 Computer Science and Engineering, University of North Texas, Denton, USA 3 Computer Information Technology, Northern Arizona University, Flagstaff, USA 4 Computer Information Technology, Northern Arizona University, Flagstaff, USA 5 Computer Information Technology, Northern Arizona University, Flagstaff, USA 6 Computer Information Technology, Northern Arizona University, Flagstaff, USA 7 Computer Science, Tunghai University, Taichung, Taiwan * Corresponding author: Tianrui Liu (Email: tianrui.liu.ml@gmail.com ) Abstract: Image captioning[1] strives to generate pertinent captions for specified images[56], situating itself at the crossroads of Computer Vision (CV) [2] and Natural Language Processing (NLP)[54]. This endeavor is of paramount importance with far- reaching applications in recommendation systems, news outlets, social media, and beyond. Particularly within the realm of news reporting, captions are expected to encompass detailed information, such as the identities of celebrities captured in the images. However, much of the existing body of work primarily centers around understanding scenes and actions.[3] In this paper, we explore the realm of image captioning specifically tailored for celebrity photographs, illustrating its broad potential for enhancing news industry practices. This exploration aims to augment automated news content generation, thereby facilitating a more nuanced dissemination of information.[57] Our endeavor shows a broader horizon, enriching the narrative in news reporting through a more intuitive image captioning framework. Keywords: Image captioning; Computer vision; Natural language processing; Content generation. 1. Introduction Image captioning is a typical topic that bridges computer vision[6] and natural language processing.[10,42] In image captioning task, we aim to generate relevant captions for given images. As a very application-oriented task, there has been a wealth of research on image captioning.[12] This technique has been widely used in many fields: traffic detection[13,24], material analysis[22], communications deployment[35,39], aerial search[37], language model designing[38], embedding development[40,41]. Meanwhile, there are much less research that considers generating captions for specific names. We believe this problem is important in the settings of news report. In such scenarios, captions generated by algorithms should take the faces into account.[4] say we intend to generate descriptions like "Obama is delivering a speech".[36] Common method may only be able to generate sentences like "a man is delivering a speech" and is unacceptable apparently.[25] In this project, we shed light on the problem of image captioning in scenarios where celebrities appear[50] and propose a combined method to solve this problem.[29] Our algorithms take three step to generate the ultimate sentences: (1) For the given image (with celebrities appear), we employ a common encoder-decoder architecture for image captioning and generate captions without names.[30] (2) We use MCTNN and Resnet[34] to get the names of faces appearing in the image. (3) We parse the output sentence in step(1) and replace the parts with celebrity names accordingly. Extensive experiments show the feasibility of our method in many simple scenarios. The structure of the report are as follows: Section 2 defines the subtask of our project, which are comprised of Image captioning, Face recognition, Noun phrase(NP) chunks matching.[51] Section 3 discusses the details of our methods and implementations. Our experiment and results are shown in Section 4. In section 5, we conclude our approach and further discuss the strength and weakness of our methods. 2. Problem Definition 2.1. Image Captioning The goal of image captioning is to generate relevant description for given images. The problem can be thought two-fold as it connects computer vision and natural language processing: 1) use an encoder architecture (CNN, Transformer)[32] to process the image; 2) use a decoder to decode the encoded image representation into sentences.[16] The model is trained in a supervised learning pattern[11] with some existing image to caption (ground truth) pairs. In the training stage we maximize the similarity of generated caption with ground truth,[14] and in the testing stage we decode the encoded image representation directly to obtain the outcome. In recent years, various efforts have been made in the image captioning tasks and the performance of caption generation has improved significantly thanks to more efficient encoder- decoder architectures.[20] 2.2. Face recognition in images Facial recognition is the task of making a positive identification of a face in a photo against a pre-existing database of faces. It begins with detection - distinguishing human faces from other objects in the image - and then works on identification of those detected faces, which can be seen as a classification problem. [15] In general, face recognition is a well-defined and relatively mature field. There has been many theoretical analysis[21] and successful engineering practice that obtain very high accuracy. In this task, we employ MTCNN to obtain the 285 bounding box for faces and use a Resnet pretrained with vggface2 to complete the classification tasks.[43] 2.3. Celebrity-aware image captioning Celebrity-aware image captioning problem combines both image captioning and face recognition. 3. Approach Our pipeline has three parts: 1. Image captioning: Image captioning modules do supervised training on datasets[49] (we use Flickr 8k/3k) and generate common (no specific names) captions in the testing stage.[7] 2. Face recognition: Face recognition employs MTCNN network to obtain bounding boxes for all appeared faces, and do classification using Inception\_v1 pretrained on vggface2.[31] 3. Noun phrase matching: Noun phrase matching module uses NLP packages and some rules to obtain noun phrase chunks (also needs to be people-related) like "a man" or "a young asian boy" and replaces them with the names derived in phrase II and generates the final outputs.[26] The overall architecture of our pipeline is shown in figure 1. Figure 1. Overall image captioning architecture 3.1. Image captioning In the image captioning step, we employ the widely-used encoder-decoder image captioning architecture. For the image side, we use a pretrained Resnet-50[53] without the last pooling and linear layer as encoder[52] and obtain a Tensor with shape [batch_size, 49, 2048]. [8]For decoder, we use a LSTM. To improve the performance of encoder-decoder, we add an Bahdanau attention layers[9] to calculate the attention between encoder outputs and initial states of the decoder.[5] The model is trained in a supervised learning setting: we first trained the parameters of LSTM to minimize the loss between generated sentences and the ground true captions. In the testing stage, we generate captions using the optimal parameters and evaluate the results by human. The architecture of the image captioning module is shown in figure 2.[27] Figure 2. Image captioning module 286 3.2. Face recognition Since our input images are scenes of behaviors instead of some simple human faces, we need to first extract face regions before performing face classification tasks. We follow the work of MTCNN [10] to complete this task: we first employ the MTCNN architecture to draw faces’ bounding boxes and then use an Inception_Resnet_v1 to obtain the name classes. The architecture of our face recognition networks are shown in picture 3.[28] Figure 3. MTCNN architecture Simply speaking, MTCNN are three cascaded convolutional networks: a Proposal Network (P-Net), a Refinement Network (R-Net) and an Output Network (O-Net). Firstly, candidate windows are produced through a fast P-Net. After that, we refine these candidates in the next stage through a R-Net. In the final stage, the O-Net produces final bounding box and facial landmarks position. 3.3. NP chunk matching Figure 4. NP chunk matching The last step is the alignment between celebrity names and the NP chunk in the generated captions. We parse the captions with NLTK and Spacy packages and obtain the people-related NP chunks. In general (sequence exchangeable) cases, sequence of the chunks are unimportant[55] so we simply sequentially align names with chunks. Using regular expression, we replace NP chunks like "A man", "The woman", "Two young boys" with celebrity names we derived in the face recognition step. Some corner cases need to be carefully coped with in this step, and in some cases it may be hard to match (or need more complicated rules).[58] For detailed implementation, please see refer to our code. In nonexchangeable cases, problems are much more challenging. We intend to align the names and faces by considering the position relation between image2text attention matrix and recognized celebrity faces. Ideally, the pixel related to faces will have large attention weight, whose position will be helpful for the alignment. However, we find that the sequence of NP chunks in the datasets[59] are usually not in accordance with the position of faces. What's more, the attention we use in out experiment are not precise enough.[48] Due to those limitations, our method is limited to solving exchangeable cases, unless we alter the architecture. 4. Experiment We use the following datasets in our project (due to stricter policy on privacy launched in recent years, we do not have access to larger datasets. Flickr 8k/30k: contains about 8,000 images collected from Flickr, together with 5 reference sentences for each image provided by human annotators. COCO Captions: contains over one and a half million captions describing over 330,000 images. For the training and validation images, five independent human generated captions are be provided for each image. We carry out supervised learning using Flickr 8k/30k or Vggface2 for captioning. Overally speaking, our pipeline can obtain a very good results with an accuracy performance over 287 90%, which demonstrates the power of combining the deep neural networks and our face recognition modules.[47] 5. Conclusion and Discussion 5.1. Discussion While our pipeline can solve the captioning problems in many cases, there are some cases we fails. The limitations of our method are as follows: Mediocre generation performance: The generation performance is not exceedingly good because we use a rather small encoder decoder architecture in our model. In addition, we employs 30,000 images together with 150,000 captions in the training process.[44] Comparing to the massive datasets used in state-of-the-art language and computer vision models, this is still a rather standard number. As our image and caption data directly decides the volume of the vocabulary dictionary, different sentences patterns and object types, the poor performance obtained by limited datasets is within our expectation. In the future, we may attempt to increase the dataset size[19] and use more powerful encoder-decoder architectures (like using pretrained Bert for the captioning part)[17] Inaccurate NP chunk matching: As we mentioned in section 4, at present our method are not capable to deal with immutable type celebrity image captioning problems due to the resolution of attention we uses and the inaccurate word sequence of Flickr datasets.[45] And in some cases, we fail to perform the matching. There are potential solutions. 1. Use more sophisticated multi-modality approach. Particularly there are much research like CLIP that proposes to connect texts with images. 2. Improve the quality of the datasets. Some datasets contains anchor boxes for each object and person in the image, which can be utilized for a more precise matching. Also, we need to carefully look into the sequence of vocabularies in the sentence. 3. Consider the entire task jointly. If we can obtain image captioning datasets containing the celebrities' name, we can customize certain loss and improve the model's ability to recognize faces. Joint model may have better grammatical accuracy then our separate-steps pipeline.[18,46] 5.2. Conclusion In this paper, we use a combined pipeline to conduct image captioning for celebrities. Our architecture uses a CNN and RNN based encoder-decoder to process the input images and generate captions. At the same time, MTCNN networks draws the bounding boxes for faces and Inception network performs classification task. Lastly, we employ NLP packages like NLTK and some engineered rules to implement NP chunk matching. [23] The incorporation of such a pipeline can significantly abbreviate the time-to-market, while ensuring a high standard of accuracy and relevance in generated content.[33] Our endeavors lay a solid groundwork, beckoning a new era of intelligent, automated news generation systems attuned to the dynamic demands of the modern digital media. References [1] Vinyals, Oriol, et al. "Show and tell: A neural image caption generator." Proceedings of the IEEE conference on computer vision and pattern recognition. 2015. [2] Ke, Lei, et al. "Reflective decoding network for image captioning." Proceedings of the IEEE/CVF international conference on computer vision. 2019. [3] Cao, Qiong, et al. "Vggface2: A dataset for recognising faces across pose and age." 2018 13th IEEE international conference on automatic face & gesture recognition (FG 2018). IEEE, 2018. [4] Zhang, Kaipeng, et al. "Joint face detection and alignment using multitask cascaded convolutional networks." IEEE signal processing letters 23.10 (2016): 1499-1503. [5] Liu, Tianrui, et al. "News recommendation with attention mechanism." arXiv preprint arXiv:2402.07422 (2024). [6] Guo, Cheng, et al. "THE ROLE OF MACHINE LEARNING IN ENHANCING COMPUTER VISION PROCESSING." АКТУАЛЬНЫЕ ВОПРОСЫ СОВРЕМЕННЫХ НАУЧНЫХ ИССЛЕДОВАНИЙ. 2023. [7] Anderson, Peter, et al. "Bottom-up and top-down attention for image captioning and visual question answering." Proceedings of the IEEE conference on computer vision and pattern recognition. 2018. [8] Zhao, Yufan, et al. "AN EXAMINATION OF TRANSFORMER: PROGRESS AND APPLICATION IN THE FIELD OF COMPUTER VISION." СОВРЕМЕННАЯ НАУКА: АКТУАЛЬНЫЕ ВОПРОСЫ, ДОСТИЖЕНИЯ И ИННОВАЦИИ. 2023. [9] Liu, Tianrui, et al. "News Recommendation with Attention Mechanism." Journal of Industrial Engineering and Applied Science 2.1 (2024): 21-26. [10] Radford, Alec, et al. "Learning transferable visual models from natural language supervision." International conference on machine learning. PMLR, 2021 [11] Li, Yanjie, et al. "Transfer-learning-based network traffic automatic generation framework." 2021 6th International Conference on Intelligent Computing and Signal Processing (ICSP). IEEE, 2021 [12] Liu, Wei, et al. "Cptr: Full transformer network for image captioning." arXiv preprint arXiv:2101.10804 (2021). [13] Liu, Tianrui, et al. "Particle Filter SLAM for Vehicle Localization." Journal of Industrial Engineering and Applied Science 2.1 (2024): 27-31. [14] Zhao, Zhiming, et al. "Enhancing E-commerce Recommendations: Unveiling Insights from Customer Reviews with BERTFusionDNN." Journal of Theory and Practice of Engineering Science 4.02 (2024): 38-44. [15] Su, Jing, et al. "Large Language Models for Forecasting and Anomaly Detection: A Systematic Literature Review." arXiv preprint arXiv:2402.10350 (2024). [16] Xiong, Jize, et al. "Decoding sentiments: Enhancing covid-19 tweet analysis through bert-rcnn fusion." Journal of Theory and Practice of Engineering Science 4.01 (2024): 86-93. [17] Liu, Shun, et al. "Financial time-series forecasting: Towards synergizing performance and interpretability within a hybrid machine learning approach." arXiv preprint arXiv:2401.00534 (2023). [18] Popokh, Leo, et al. "IllumiCore: Optimization Modeling and Implementation for Efficient VNF Placement." 2021 International Conference on Software, Telecommunications and Computer Networks (SoftCOM). IEEE, 2021. [19] Su, Jing, Suku Nair, and Leo Popokh. "EdgeGYM: a reinforcement learning environment for constraint-aware NFV resource allocation." 2023 IEEE 2nd International Conference on AI in Cybersecurity (ICAIC). IEEE, 2023. 288 [20] Su, Jing, Suku Nair, and Leo Popokh. "Optimal resource allocation in sdn/nfv-enabled networks via deep reinforcement learning." 2022 IEEE Ninth International Conference on Communications and Networking (ComNet). IEEE, 2022. [21] Fu, Zhe, Xi Niu, and Li Yu. "Wisdom of crowds and fine- grained learning for serendipity recommendations." Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval. 2023. [22] Wei, Xiaojun, et al. "Narrowing Signal Distribution by Adamantane Derivatization for Amino Acid Identification Using an α-Hemolysin Nanopore." Nano Letters (2024). [23] Song, Ge, et al. "Energy consumption auditing based on a generative adversarial network for anomaly detection of robotic manipulators." Future Generation Computer Systems 149 (2023): 376-389. [24] Ou, Junlin, et al. "Hybrid path planning based on adaptive visibility graph initialization and edge computing for mobile robots." Engineering Applications of Artificial Intelligence 126 (2023): 107110. [25] Zhou, Yucheng, et al. "Visual In-Context Learning for Large Vision-Language Models." arXiv preprint arXiv: 2402.11574 (2024). [26] Zhou, Yucheng, et al. "Thread of thought unraveling chaotic contexts." arXiv preprint arXiv:2311.08734 (2023). [27] Zhou, Yucheng, et al. "Claret: Pre-training a correlation-aware context-to-event transformer for event-centric generation and classification." arXiv preprint arXiv:2203.02225 (2022). [28] Tian, Jiwei, et al. "LESSON: Multi-Label Adversarial False Data Injection Attack for Deep Learning Locational Detection." IEEE Transactions on Dependable and Secure Computing (2024). [29] Tian, Jiwei, et al. "Adversarial attacks and defenses for deep- learning-based unmanned aerial vehicles." IEEE Internet of Things Journal 9.22 (2021): 22399-22409. [30] Wu, Jing, et al. "SwitchTab: Switched Autoencoders Are Effective Tabular Learners." arXiv preprint arXiv:2401.02013 (2024). [31] Chen, Suiyao, et al. "Recontab: Regularized contrastive representation learning for tabular data." arXiv preprint arXiv:2310.18541 (2023). [32] Chen, Suiyao, et al. "Claims data-driven modeling of hospital time-to-readmission risk with latent heterogeneity." Health care management science 22 (2019): 156-179. [33] Chen, Suiyao, et al. "A data heterogeneity modeling and quantification approach for field pre-assessment of chloride- induced corrosion in aging infrastructures." Reliability Engineering & System Safety 171 (2018): 123-135. [34] Hsieh, Yung-Ting, Khizar Anjum, and Dario Pompili. "Ultra- low Power Analog Recurrent Neural Network Design Approximation for Wireless Health Monitoring." 2022 IEEE 19th International Conference on Mobile Ad Hoc and Smart Systems (MASS). IEEE, 2022. [35] Hsieh, Yung-Ting, Zhuoran Qi, and Dario Pompili. "ML-based joint Doppler estimation and compensation in underwater acoustic communications." Proceedings of the 16th International Conference on Underwater Networks & Systems. 2022. [36] Sun, Chuanneng, et al. "Fed2kd: Heterogeneous federated learning for pandemic risk assessment via two-way knowledge distillation." 2022 17th Wireless On-Demand Network Systems and Services Conference (WONS). IEEE, 2022. [37] Sun, Chuanneng, Songjun Huang, and Dario Pompili. "HMAAC: Hierarchical Multi-Agent Actor-Critic for Aerial Search with Explicit Coordination Modeling." 2023 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2023. [38] Sun, Chuanneng, et al. "Contextual Biasing of Named-Entities with Large Language Models." arXiv preprint arXiv: 2309.00723 (2023). [39] Younis, Ayman, Chuanneng Sun, and Dario Pompili. "Communication-efficient Federated Learning Design with Fronthaul Awareness in NG-RANs." 2022 IEEE 19th International Conference on Mobile Ad Hoc and Smart Systems (MASS). IEEE, 2022. [40] Flynn, Patrick, Xinyao Yi, and Yonghong Yan. "Exploring source-to-source compiler transformation of OpenMP SIMD constructs for Intel AVX and Arm SVE vector architectures." Proceedings of the Thirteenth International Workshop on Programming Models and Applications for Multicores and Manycores. 2022. [41] Yi, Xinyao, et al. "CUDAMicroBench: Microbenchmarks to Assist CUDA Performance Programming." 2021 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW). IEEE, 2021. [42] Li, Xiaying, Belle Li, and Su-Je Cho. "Empowering Chinese Language Learners from Low-Income Families to Improve Their Chinese Writing with ChatGPT’s Assistance Afterschool." Languages 8.4 (2023): 238. [43] Xie, Ying, et al. "Advancing Legal Citation Text Classification A Conv1D-Based Approach for Multi-Class Classification." Journal of Theory and Practice of Engineering Science 4.02 (2024): 15-22. [44] Luo, Yang, et al. "Enhancing E-commerce Chatbots with Falcon-7B and 16-bit Full Quantization." Journal of Theory and Practice of Engineering Science 4.02 (2024): 52-57. [45] Pan, Zhenyu, et al. "Ising-traffic: Using ising machine learning to predict traffic congestion under uncertainty." Proceedings of the AAAI Conference on Artificial Intelligence. Vol. 37. No. 8. 2023. [46] Pan, Zhenyu, et al. "CoRMF: Criticality-Ordered Recurrent Mean Field Ising Solver." arXiv preprint arXiv:2403.03391 (2024). [47] Hu, Zhirui, et al. "On the design of quantum graph convolutional neural network in the nisq-era and beyond." 2022 IEEE 40th International Conference on Computer Design (ICCD). IEEE, 2022. [48] Chen, Yinda, et al. "Self-supervised neuron segmentation with multi-agent reinforcement learning." arXiv preprint arXiv:2310.04148 (2023). [49] Chen, Yinda, et al. "Learning multiscale consistency for self- supervised electron microscopy instance segmentation." arXiv preprint arXiv:2308.09917 (2023). [50] Chen, Yinda, et al. "Generative text-guided 3d vision-language pretraining for unified medical image segmentation." arXiv preprint arXiv:2306.04811 (2023). [51] Tong, Xin, et al. "A Deep‐Learning Approach for Low‐Spatial‐ Coherence Imaging in Computer‐Generated Holography." Advanced Photonics Research 4.1 (2023): 2200264. [52] Xu, Renjun, et al. "$ E (2) $-Equivariant Vision Transformer." Uncertainty in Artificial Intelligence. PMLR, 2023. [53] Gao, Shangde, et al. "Contrastive Knowledge Amalgamation for Unsupervised Image Classification." International Conference on Artificial Neural Networks. Cham: Springer Nature Switzerland, 2023. [54] Shangguan, Zhongkai, Zihe Zheng, and Lei Lin. "Trend and thoughts: Understanding climate change concern using 289 machine learning and social media data." arXiv preprint arXiv:2111.14929 (2021). [55] Shangguan, Zhongkai, et al. "Neural process for black-box model optimization under bayesian framework." arXiv preprint arXiv:2104.02487 (2021). [56] Zang, Hengyi. "Precision calibration of industrial 3d scanners: An ai-enhanced approach for improved measurement accuracy." Global Academic Frontiers 2.1 (2024): 27-37. [57] Wang, Yishuang, et al. "Predicting Nonlinear Structural Dynamic Response of ODE Systems Using Constrained Gaussian Process Regression." Society for Experimental Mechanics Annual Conference and Exposition. Cham: Springer Nature Switzerland, 2023. [58] Platz, Roland, Xinyue Xu, and Sez Atamturktur. "Introducing a Round-Robin Challenge to Quantify Model Form Uncertainty in Passive and Active Vibration Isolation." Society for Experimental Mechanics Annual Conference and Exposition. Cham: Springer Nature Switzerland, 2023. [59] Li, Zhenglin, et al. "Comprehensive evaluation of Mal-API- 2019 dataset by machine learning in malware detection." International Journal of Computer Science and Information Technology 2.1 (2024): 1-9.