id	author	title	date	pages	extension	mime	words	sentence	flesch	summary	cache	txt
fuelectenerg-13355	KardanMoghaddam, Hossein; Akbarimajd, Adel; Ranjbarpour, Mohammad; Nooshyar, Mahdi; Jamali, Shahram	A STRUCTURE BASED ON TROCR TRANSFORMER AND LARGE LANGUAGE MODEL FOR CLASSIFICATION OF HANDWRITTEN TEXTS	2025	18	.pdf	application/pdf	9030	482	51	The proposed method in this research involves extracting handwritten texts from images within a database dataset and converting them into text data. Table 1 Pre-Trained Models TrOCR Pre-trained Models Features Microsoft/TrOCR- base-handwritten Based on “Base” Transformer structure especially has been designed for handwritten text and optimized for recognition of text in images containing handwritten texts Microsoft/TrOCR- small-handwritten The smaller version of TrOCR(Small)with less computational amount is suitable for systems with limited computational resources with acceptable performance in handwritten recognition Microsoft/TrOCR- large-handwritten The large version with high capacity for better learning of convoluted patterns, is suitable for cumbersome tasks requiring high precision like processing sensitive or complicated documents, suitable for use in research projects, reading handwritten texts with complicated details and organizational and industrial usages requiring high precision These pre-trained models by Microsoft have been trained using general OCR data and handwritten texts and after that the obtained information has been saved on a CSV file then this CSV file contains obtained texts from images is given to LLM as an input so that the extracted texts are categorized into different subjects (“Health”, “Technology”, “Finance”, “Education”, “Sports”, “Others”) based on the content.	cache/fuelectenerg-13355.pdf	txt/fuelectenerg-13355.txt
