paper title (use style: paper title) issn : vol.1 no.1, 01 january 2020 buana information tchnology and computer sciences (bit and cs) 16 | vol.1 no.1, 01 january 2020 buana information tchnology and computer sciences (bit and cs) designing and building client-server based student admission applications arip solehudin 1 teknik informatika school of computer science universitas singaperbangsa karawang arip.solehudin@gmail.com nono heryana 2 teknik informatika school of computer science universitas singaperbangsa karawang nonoheryana@staff.unsika.ac.id ‹β› yana cahyana 3 teknik informatika school of engineering and computer science universitas buana perjuangan karawang yana.cahyana@ubpkarawang.ac.id abstract—admission of new students is a routine activity organized by the educational institution. early childhood education (paud) is included in out-of-school education in the age range of two to five years, the goal is to help improve physical and spiritual growth and development so that children have the readiness to enter further education. admission of new students to al-qudwah is still done using prospective students visiting al-qudwah and filling out the registration form with paper. new student admission application at al quwah aims to promote the school to the community at large and can also be registered without parents of prospective students visiting educational institutions. this application makes it easy for parents of prospective students to find out the facilities and infrastructure and the information available on al quwah. we built this application using html and phpprogramming languages with the php mysql database and using the waterfall method. keywords—application, admission, website,uml, waterfall i. introduction paud-level education is currently in demand by parents in helping the growth and development of children with the hope that children are better prepared to enter the next level of education. before entering, of course, parents will consider and look for information about which paud they will choose to assume responsibility for educating their children. usually, the information that parents look for about the condition of their school, the curriculum taught, admission fees, monthly administration, and what learning activities are available at these paud institutions. the development of information technology currently plays an important role in the development of information, especially internet users because the results of the information technology make communication limited by time and space to make communication without space and time restrictions. with the presence of the internet to provide alternatives for users in utilizing the information search needed. through the internet, it can provide the speed of information every time requires information, detail and free of cost. so that in indonesia there are many internet users every year. in general, the purpose of using the internet is to obtain and share information. the development of the internet has penetrated the field of education to make competitive competition between educational institutions. ii. theoretical basis a. understanding new students acceptance is a welcome process, act or attitude towards someone. students are students at an academy or college. new is something that did not exist before. new student registration (psb) is the activity of accepting and selecting prospective participants in education and training at schools. admission of new students (psb) is a process of academic selection of prospective students towards higher education [3]. b. website the website is one application that contains multimedia documents (text, images, sound, animation, video) in it that uses the http protocol (hypertext transfer protocol) and to access it using a software called a browser. some types of popular browsers today include internet explorer produced by microsoft, mozilla firefox, opera, and safari produced by apple [2]. c. uml unified modeling language (uml) is a standard specification language used to document, specify and build software. uml is a methodology in developing objectoriented systems and is also a tool to support system development [4]. d. waterfall the method used by the author in developing this software uses the waterfall model. the waterfall model is often also called the sequential linear or classic life cycle model. the waterfall model provides a sequential or sequential software life cycle approach starting from analysis, design, coding, testing and support stages [5]. e. system is working mobility of procedures or elements that are interconnected, gathered together to carry out an activity or to solve a problem [1]. iii. research methods the method of developing applications for new student admissions was created using the waterfall method. 17 | vol.1 no.1, 01 january 2020 buana information tchnology and computer sciences (bit and cs) fig. 1. waterfall method iv. result and discussion a. proposed application design 1. use case diagrams fig. 2. use case operator diagram 2. activity diagram activity diagram illustrates the workflow or work activities of the application to be implemented. the following is an activity diagram of the proposed application that will be implemented. fig. 3. activity diagram 3. class diagram class diagrams are used to describe the relationship between the classes of a system. following is the class diagram of the new student admissions application. fig. 4. class diagram b. implementation 1. main page display 18 | vol.1 no.1, 01 january 2020 buana information tchnology and computer sciences (bit and cs) fig. 5. main page display 1. psb display fig. 6. psb display conclusion the conclusions the author can draw that are: the acceptance system in al qudwah paud after observation is still manual so it is feared the loss of registration data because of damage or fire, with the website application developed by the author can minimize losses data because registration data is stored on a database. new student admission applications can media for publication and also to develop information technology, so that information got is fast and efficient. storage of registration records is still done in physical form, in this application each participant has a personal account to view the track record and results got. references [1] bayu priyatna, penerapan metode user centered design (ucd) pada sistem pemesanan menu kuliner nusantara berbasis mobile android, account. inf. syst. penerapan, pp. 17–30, 2018. [2] arief, m. rudyanto., 2011. pemrograman web dinamis menggunakan php dan mysql: andi publisher. [3] hendini, ade., 2016. pemodelan uml sistem informasi monitoring penjualan dan stok barang (studi kasus: distro zhezha pontianak). jurnal khatulistiwa informatika, vol. iv, no. 2 desember 2016, hlm. 107-116. [4] witanto, regi., solihin, hanhan hanafiah.,2016. perancangan sistem informasi penerimaan siswa baru berbasis web (studi kasus : smp plus babussalam bandung). jurnal infotronik volume 1, no. 1, desember 2016, hlm. 54-63. [5] pranata, dana., & hamdani., & k, marisa, dyna. 2015. rancang bangun website jurnal ilmiah bidang komputer. jurnal informatika mulawarman.vol.10. paper title (use style: paper title) issn : vol.1 no.1, 01 january 2020 buana information tchnology and computer sciences (bit and cs) 1 | vol.1 no.1, 01 january 2020 buana information tchnology and computer sciences (bit and cs) application of prototype method on student monitoring system based on web april lia hananto 1 information systems study program school of engineering and computer science universitas buana perjuangan karawang aprilia.hananto@ubpkarawang.ac.id bayu priyatna 2 information systems study program school of engineering and computer science universitas buana perjuangan karawang bayu.priyatna@ubpkarawang.ac.id ‹β› asep haris 3 information systems study program school of engineering and computer science universitas buana perjuangan karawang asepharis88@gmail.com abstract—public vocational school in subang which continues to improve its academic activities, specifically in terms of improving student discipline related to student participation in school and improving student achievement. the collection of information about student participation and the value of new students delivered at the end of the semester compilation of report cards makes students who experience difficulties in the development of grades, meetings and student learning activities at school, for that we need a system that can help students to help guardians of students and the school in activities that involve students in school. in this research, we use a prototype method for system development. the advantage of the system built is that it can send attendance messages or student grades to enter no answers or get poor test scores for student guardians, student guardians can provide feedback on incoming information and make permits through the existing information system pages, the school also can use attendance data and grades to support activities at school. keywords— a student monitoring system, attendance, prototype. i. introduction monitoring academy activities is a major activity in the world of education. cipunagara vocational high school 1 as one of the educational institutions, of course, must carry out these activities as mandatory activities in the implementation of student learning activities, but these activities cannot be carried out optimally. one of the problems that often occur in high school level education environments is monitoring student attendance, the process of delivering information from school to student guardians in addition to internal problems such as the flow of data that is processed quite a lot every day, there are also problems caused by external factors such as manipulation data or cheating done by students in terms of attendance, this must be addressed immediately because it is very disruptive to the learning process at school. also, guardians of students still experience difficulties in monitoring the progress of their children's learning achievements or activities at school. notification of achievements (grades) and abscesses is done when the school report card is dropped, which is once every end of the semester. student guardians can only get the final results of their children's learning activities without being able to monitor their children's achievements and attendance during the teaching and learning process in progress. notification for students with problems is done by sending a letter and sometimes the letter is not delivered. also, the assessment of the teacher is also still not computerized, so the guardians of students or students still have difficulty knowing the value during the teaching and learning process. based on the above background, we need an information system that can facilitate teachers in absenteeism, provide assessments to students and facilitate student guardians in monitoring their children's academic activities. ii. theoretical basis a. prototyping method the prototyping method is an iterative process in the development of systems where requirements are changed into working systems that are continuously improved through collaboration between users and analysts". they can build prototyping methods through several development tools to simplify the process and the following cycle prototype methods [7]. fig. 1. the prototype method cycle [7]. the making of a prototype for system developers aims to collect information from users so that users can interact with the prototype models that are developed because prototypes describe the initial version of the system for the continuation of the actual larger system [1]. b. application i can interpret applications as a software program that runs on a particular system useful to assist various activities carried out by humans [2]. to understand the application is a program ready to use that is made to carry out a function for 2 | vol.1 no.1, 01 january 2020 buana information tchnology and computer sciences (bit and cs) users of application services and a target to be addressed can use the use of other applicationsa target to be addressed can use the use of other applications. according to the executive computer dictionary, the application has a meaning that is problem solving that uses one of the application's data processing techniques which refers to a desired or expected computation or the expected data processing [3] c. system the system is a network of procedures that are interconnected procedures, gathered together to carry out an activity or to complete a particular goal [4]. a sistem is a group of elements that are integrated with the common porpose of achieving an objective [5]. the system is a network of interrelated procedures, gathered together to carry out an activity or for a particular purpose [6]. iii. research methods in this study, applying the prototype development method to design and implement system designs. to obtain maximum results, this study therefore emphasizes the needs analysis. iv. result and discussion a. overview of the proposed system the physical architecture of the system comprises three main parts, client, application server, and database server. i can see the working principle of the system as a whole in the following figure: admin smartphone selver simson db guru/wali kelas wali siswa wali wali siswa wali siswa router wifi acess point firewall printer laptop switch fig. 2. architecture of student monitoring information systems b. use case diagram in the n use case of the design of student monitoring information systems, where there are 4 actors admin and homeroom teacher, subject teacher and student guardian. where 1. admin who has full rights in the activities of input, edit and delete, 2. homeroom teacher may see attendance data / grades, print attendance data / values and view and input messages or information for teachers or students, 3. subject teachers may input and view attendance data or grades, view incoming permits, view and input information either to the student, or two students, and 4. guardians students may view data on grades, attendance, information, input permits, and input responses incoming information. guru mata pelajaran wali kelas admin wali siswa pusat informasi input absensi / nilai lihat data absensi / nilai pengajuan ijin cetak data absensi/nilai mengelola master data kirim pesan login <> << in cl ud e> > <> <> <> <> fig. 3. use case diagram student monitoring information system c. class diagram after the use case diagram, the writer makes a class diagram aimed at a container that describes the structure of objects in the system formed from the relationships between classes. fig. 4. class diagram of student monitoring information system d. implementation of the user interface implementation of the user interface is done with every display program that is built. the following is the implementation of the user interface application simson (online student monitoring information system) created. e. main page of the program this page is the first page that appears when opening the simson application (online student monitoring information system), this page contains the logo, the name of the school agency and a portal to log in as the guardian of students and the school (subject teacher, homeroom teacher and admin); 3 | vol.1 no.1, 01 january 2020 buana information tchnology and computer sciences (bit and cs) fig. 5. main page of the program f. attendance data input page this student attendance page is the page that appears after the teacher fills in the attendance form, and clicks the show button, in this form the teacher can do student attendance, the data will be displayed in the default settings present, so the teacher can change the attendance status of students who are not present course, the following appearance of the user interface page design attendance (student attendance); fig. 6. attendance data input page g. value monitoring this value monitoring page is a page the student guardians can use that to see the values gottheir children can obtain that per subject based on the selected semester and class display the implementation of the value monitoring page: fig. 7. value monitoring h. print page value data this value data print page is a page that the teacher can use if he wants to print the value data that has been displayed. display the implementation of the value data report: fig. 8. print page value data i. attendance data print page display the attendance report data: 4 | vol.1 no.1, 01 january 2020 buana information tchnology and computer sciences (bit and cs) fig. 9. attendance data print page conclusion based on the results of the analysis conducted by the results of the design, realization, and testing of the system, several conclusions can be drawn, including: the results of the system analysis that runs include attendance data, student grades, and notification of problematic student information to student guardians, processed to get a new system design that can overcome the problems expressed in the background.the results of the new system design are discussed in the form of a web-based and mobile student monitoring system that can manage student attendance data, grades and infringement information so that it can be accessed by student guardians computerized or mobile. references [1] d. purnomo, “model prototyping pada pengembangan sistem informasi,” jimp j. inform. merdeka pasuruan, vol. 2, no. 2, pp. 54–61, 2017. [2] b. p. baenil huda, “penggunaan aplikasi content manajement system (cms) untuk pengembangan bisnis berbasis e-commerce,” systematics, vol. 1, no. 2, pp. 81–88, 2013. [3] andi juansyah, “pembangunan aplikasi child tracker berbasis assisted – global positioning system ( a-gps ) dengan platform android,” j. ilm. komput. dan inform., vol. 1, no. 1, pp. 1–8, 2015. [4] bayu priyatna, “penerapan metode user centered design (ucd) pada sistem pemesanan menu kuliner nusantara berbasis mobile android,” account. inf. syst. penerapan, pp. 17–30, 2018. [5] rini, “sistem informasi pengolahan data penanggulangan bencana pada kantor badan penanggulangan bencana daerah (bpbd) kabupaten padang pariaman,” vol. 3, no. june, 2016. [6] ermatita, “analisis dan perancangan sistem informasi perpustakaan,” j. sist. inf., vol. 8, no. 1, pp. 966–977, 2016. [7] muharto & ambarita arisandy. 2016. metode penelitian sistem informasi. yogyakarta: deepublish. paper title (use style: paper title) issn : vol.1 no.1, 01 january 2020 buana information tchnology and computer sciences (bit and cs) 5 | vol.1 no.1, 01 january 2020 buana information tchnology and computer sciences (bit and cs) academic application design web-based on junior high schools baenil huda 1 information systems study program school of engineering and computer science universitas buana perjuangan karawang baenilhuda@ubpkarawang.ac.id fitri nurapriani 2 information systems study program school of engineering and computer science universitas buana perjuangan karawang fitri.apriani@ubpkarawang.ac.id ‹β› helga amanda3 information systems study program school of engineering and computer science universitas buana perjuangan karawang si15.helgaamanda@mhs.ubpkarawang.ac.id abstract—current advances in information technology have provided great benefits in the world of education, making webbased academic applications is a major use of information technology. information technology enables academic data to be processed and, making the required presentation of academic information be got, and. this research uses technological trends in managing academic administration so that conventional bookkeeping in junior high schools is overcome by computer systems. the method in developing the system uses a waterfall with web-based device implementation. the application of this new system can improve the knowledge and skills of employees, teachers, and principals in web-based academic applications. keywords— application, academic, web based, waterfall. i. introduction the development of information technology requires accuracy in data processing for all institutions or institutions. the more precise in data processing, the easier it will be to get the trust of consumers, so that each institution or institution uses a structured information system that can answer and process data, including school academic data processing. using information in schools, besides improving quality, will also facilitate the administration such as business governance of existing data such as student data, teacher data, class data, student learning outcomes data, and subjects. data management like this will be easier and more effective by using academic information systems. junior high schools one purwasari in processing academic data processing student data, teacher data, lesson schedules, and others are still using a manual system. with a total class of 27 classes, comprising class vii (9 classes), viii (8 classes), ix (8 classes). the number of students in each class is almost 40 students and the number of teachers is almost 30, so it is very inefficient in processing academic data. by looking at the fun challenging in junior high schools one purwasari, the author wants to create an academic information system and take the title "web-based academic application design at junior high schools one purwasari". ii. theoretical basis a. design the design is a visual form that results from the planned creative forms. design is the process of planning everything [1]. system design is a phase in which design expertise is needed for computer elements that will use the system, namely the selection of equipment and computer programs for the new system [2]. b. application application is a ready-made program that can run commands from the user of the application to get more accurate results following the purpose of making the application. application is a problem solving that uses one of the application data processing techniques which refers to a desired or expected computation or expected data processing [3]. c. academic academics are a condition where people can convey and accept ideas, thoughts, science, and be able to test them [4], the word academic comes from the greek academic which means a public park (plasma) in the northwest of athens. after that, the word academic turned into academic, which is a kind of college place [5]. d. website a website is a collection of pages where a page is linked to other pages [6]. website is a collection of web pages that are interconnected and it relates the files to one another [7]. iii. research methods in this study refers to the waterfall model. at each stage in this study does not always depend on the user, only an analysis phase is a step approach to the user. therefore, in developing the software, researchers used the waterfall method. iv. result and discussion a. analysis of the current system the process of data processing and academic information storage at junior high scholes one purwasari karawang still uses a manual system, it requires quite a long time to make the process of reporting information. the following is a flow map system that runs in data processing: 6 | vol.1 no.1, 01 january 2020 buana information tchnology and computer sciences (bit and cs) fig 1. the system flowmap is running b. analysis of the proposed system the proposed system analysis is based on the website using the php programming language and using the mysql database. the use of the website can streamline time and make it easier for staff and teacher sections in processing data such as student data, teacher data, class schedule data, and grade data. besides the school principal can easily access information from the website. here is the proposed academic information system flow map: staff guru kepala sekolah input data pelajaran, data guru, data siswa, data kelas data tersimpan mengambil data siswa laporan / report data input data nilai siswa data tersimpan laporan / report data database ya tidak ya tidak fig 2. activity diagram of the proposed system c. class diagram class diagrams are used fatherly to illustrate the relationship between classes of a system. the following class diagram of the academic information system application. guru #kode_guru +nip +nama_guru +jenis_kelamin +alamat +no_telp +status_aktif +add() +edit() +delete() siswa #kode_siswa +nis +nama_siswa +jenis_kelamin +agama +tempat lahir +tgl_lahir +alamat +no_telp +foto +thn_angkatan +status +add() +edit() +delete() kelas #kode_kelas +thn_ajar +kelas +nama_kelas +kode_guru +status_aktif +add() +edit() +delete() nilai #id +semester +kode_pelajaran +kode_guru +kode_kelas +kode_siswa +nilai_tugas1 +nilai_tugas2 +nilai_uts +nilai_uas +keterangan +add() +edit() +delete() pelajaran #kode_pelajaran +nama_pelajaran +keterangan +add() +edit() +delete() kelas_siswa #id #kode_kelas #kode_siswa +add() +edit() +delete() n+ 1 1 n+ 1 n+ 1 n+ n+ n+ n+ n+ 1 n+ n+ fig 3. class diagram academic information system d. application testing at the application testing stage in this thesis is done using the black-box testing method, wherein this testing is carried out functional testing from the user's side. e. maintenance maintenance is carried out to maintain applications that have been implemented. this implementation is planned periodically. it is expected that the application can continue to run also can be known as the potential for unwanted things such as data theft. conclusion the conclusion that can be drawn from the making of information systems academics at purwasari one junior high school are as follows; this web-based academic information system can assist in the processing and archiving of academic data, namely; student data, teacher data, lesson data, class data, lesson schedules, and student grades. staff and teacher data management, student subject management reports, and student grade management reports quickly. references [1] w. hidayat, “perancangan media video desain interior sebagai salah satu penunjang promosi dan informasi di pt . wans desain group,” cerita, vol. 2, no. 1, pp. 35–49, 2016. [2] p. kristanto, “ekologi industri,” j. chem. inf. model., vol. 5, no. 2011, pp. 1689–1699, 2013. [3] h. abdurahman and a. r. riswaya, “aplikasi pinjaman pembayaran secara kredit pada bank yudha bhakti,” j. comput. bisnis, vol. 8, no. 2, pp. 61–69, 2014. [4] d. m. farid suryandani, basori, “pengembangan sistem informasi akademik berbasis web sebagai sistem pengolahan nilai siswa di smk negeri 1 kudus,” j. chem. inf. model., vol. 53, no. 9, pp. 1689–1699, 2013. [5] s. k. dewi anggun kumalasari, hasan nur faozi, “perancangan sistem informasi administrasi 7 | vol.1 no.1, 01 january 2020 buana information tchnology and computer sciences (bit and cs) keuangan sekolah berbasis multiuser pada madrasah tsanawiyahal uswah bergas,” 입법학연구, vol. 제13집 1호, no. may, pp. 31–48, 2016. [6] a. d. anggi s, eko r, “perancangan sistem informasi berbasis website subsistem guru di sekolah pesantern islam 99 rancabango,” sttgarut, 2015. [7] a. suhadya, “perancangan website sebagai media promosi,” j. inform. pelita nusant., vol. 3, no. 1, pp. 82–86, 2013. paper title (use style: paper title) issn : vol.1 no.1, 01 january 2020 buana information tchnology and computer sciences (bit and cs) 12 | vol.1 no.1, 01 january 2020 buana information tchnology and computer sciences (bit and cs) teacher monitoring application in teaching based on codeigniter framework in high schools syahri susanto 1 school of oil technic akademi minyask dan gas balongan syahri28@gmail.com bayu priyatna 2 information systems study program school of engineering and computer science universitas buana perjuangan karawang bayu.priyatna@ubpkarawang.ac.id ‹β› fakhreza aditya permana 3 information systems study program school of engineering and computer science universitas buana perjuangan karawang aditya.fakhreza@gmail.com abstract—this is teacher monitoring in teaching at one karawang middle school, is needed to monitor teachers in teaching in class so it can be seen whether the teacher often does teaching and learning activities or not, teachers do not always attend school then carry out teaching and learning activities and also not teachers are not in attendance does not give assignments as a substitute for learning. this monitoring application can support the teacher monitoring process in teaching in the classroom. the system development method used is a waterfall and uses php and mysql. this application is a consideration of how often the teacher teaches in class. keywords—monitoring application, teaching teachers, php and mysql. i. introduction i need monitoring to get facts, data, and information from an activity that has been carried out. this monitoring intends to see whether the teacher's performance in carrying out teaching activities is going well so i can see it how often each teacher in carrying out teaching activities in class. monitoring the curriculum work unit, monitoring the teaching activities of teachers in the classroom. monitoring by involving the picket teacher as a direct observer of teaching activities carried out by the teacher so that, the picket teacher gives an indicator value to each teacher under the existing indicator criteria. i will take the results of the monitoring into consideration for the next semester's curriculum for the provision of teaching hours to each teacher after seeing the teacher's activeness in teaching in the classroom in the current semester. this research produces a website-based application that can help in the monitoring process that is carried out through the website. ii. theoretical basis a. software engineering software engineering is an engineering discipline or engineering that deals with all aspects of making software that requires following a systematic and approach and using appropriate tools and techniques under the problem to be solved, development constraints and resources [1]. b. monitoring to get an implementation plan is following what is planned management must prepare a program that is monitoring, monitoring will be aimed at getting facts, data and information about implementing the program, whether carrying out activities following what has been planned. the findings of the monitoring results are information for the evaluation process so that the results are whether the programs that are established and implemented get the results or not. monitoring is an activity to find out i made whether the program that is going well as it should be as planned, are there obstacles that occur and how the program implementers overcome these obstacles [2]. c. website website is one application that contains multimedia documents (text, images, sound, animation, video) in it that uses the http protocol (hypertext transfer protocol) and to access it using a software called a browser [3]. d. unified modeling language (uml) uml (unified modeling language) is a modeling language for systems or software that is an object-oriented paradigm. modeling (modeling) in analyzing and discussing a database, uml (unified modeling language) diagrams can be used. uml is one of the modeling tools to complete object-oriented software modeling [4]. iii. research methods metode waterfall software waterfall method in this research. the study used the waterfall system development method. where the method has 5 stages in developing the system. the following are the stages of system development in research. 1. design the design is used to make an initial description of the application to be made, in system modeling, web-based 13 | vol.1 no.1, 01 january 2020 buana information tchnology and computer sciences (bit and cs) teacher monitoring application interface at high scchool 1 karawang. 2. making the program code at this stage, the authors do the application coding following the design that has been made. 3. testing they carry the test out to test both and of the application that has been made, to ensure the application is running. in this study, testing using the black box testing method. 4. maintenance i do maintenance after the application runs by checking the application running. iv. result and discussion a. use case proposed application diagram in this diagram will discuss what is in this application and who can use it. 1. use case diagram application proposal for access rights to the monitoring and evaluation section. fig 2. use case diagram application proposals for access rights to the monitoring and evaluation section b. activity diagram application proposed activity diagram illustrates the workflow or work activities of the application to be implemented. the following is an activity diagram of the proposed application that will be implemented. 1. activity diagram application proposal for access rights to the monitoring and evaluation section; fig 3. activity diagram application proposal for access rights to the monitoring and evaluation section 2. activity diagram application proposal for teacher picket access rights. fig 4. activity diagram application proposal for access rights to the monitoring and evaluation section c. class diagram application proposal class diagrams are used to describe the relationship between classes of a system. the following class diagram of the monitoring application. fig 5. class diagram 14 | vol.1 no.1, 01 january 2020 buana information tchnology and computer sciences (bit and cs) d. proposed application interface design 1. login interface design to create interaction between user and application in login, the login interface design will later interact with the login process between user and application. fig 6. login interface design 2. the interface design of teacher picket access rights fig 7. interface design of teacher picket access rights 3. the interface design of m&e section access rights fig 8. interface design of m&e e. implementation this implementation contains the results of the design that has been created. the following is the display application. 2. interface the main page fig 9. main page conclusion they base the conclusion on the analysis that has been carried out as follows. 1. the teacher monitoring system in teaching at the karawang high school one work unit is still manual, in inputting indicators that are still manual so they are prone to errors. 2. in the monitoring application proposed by the author, a monitoring instrument input form is provided which is stored into the database and data from the inputted values can be seen. 3. the process of the report can be taken following the needs of the report date. references [1] b. p. april lia hananto, “rancang bangun aplikasi informasi harga produk,” technoxplore, vol. 2, no. 1, pp. 10–20, 2017. [2] arisantoso., sanwasih, moch., samsudin, didin adhuri. 2014. prototipe monitoring dan evaluasi kinerja dosen untuk menunjang kegiatan tridharma perguruan tinggi (studi kasus universitas islam attahiriyah). teknologi informasi dan multimedia. issn: 2302-3805. [3] handayani, sutri. 2018. perancangan sistem informasi penjualan berbasis e-commerce studi kasus toko kun jakarta. ilkom jurnal ilmiah. vol. 10. e-issn 2548-7779. [4] pranata, dana., & hamdani., & k, marisa, dyna. 2015. rancang bangun website jurnal ilmiah bidang komputer (studi kasus : program studi ilmu komputer universitas mulawarman). jurnal informatika mulawarman.vol.10. [5] riyalda, bondan f., & turyana iyan., & santoso eko w. 2018. sistem informasi bencana tanah longsor (sibenar) berbasis web kecamatan cililin, kabupaten bandung barat. jurnal alami. vol. 2. eissn: 2548-8635. [6] suhartanto, medi. 2012. pembuatan website sekolah menengah pertama negeri 3 delanggu dengan 15 | vol.1 no.1, 01 january 2020 buana information tchnology and computer sciences (bit and cs) menggunakan php dan mysql. sentra pendidikan engineering dan edukasi. vol. 4. issn: 2088-0154. paper title (use style: paper title) issn : 2715-2448 | e-issn : 2715-7199 vol.2 no.1 january 2021 buana information technology and computer sciences (bit and cs) 17 | vol.2 no.1, january 2021 applying the prototype model into the electronic reporting system for the elementary school student base on android tukino 1 information system, faculty of engineering and computer science universitas buana perjuangan karawang,indonesia tukino@ubpkarawang.ac.id siti masruroh 2 information system, faculty of engineering and computer science universitas buana perjuangan karawang,indonesia sitimasruroh@ubpkarawang.ac.id ‹β› daryanto herdiana 3 information system, faculty of engineering and computer science universitas buana perjuangan karawang,indonesi daryantoherdiana@gmail.com abstract— teaching and learning is an activity that is bound by goal directed and carried out specifically to achieve that goal. because it is very important to seek knowledge for a bright future. supervision of students by the guardians of the students made the results of their children's achievements not improving. as well as student assessment by the teacher is still not well managed because it is still in the form of a note report. the system method used is the prototype model. with observation and direct interviews with the student section regarding the assessment system in the school where the author researched. the results of this study are applications that can be operated on an android smartphone. this application can provide fast information and update automatically in obtaining information on student learning outcomes. keywords: student assessment, e-report card, prototype method, application. abstrak— proses belajar mengajar merupakan suatu kegiatan yang terikat dengan tujuan yang diarahkan dan dilaksanakan secara khusus untuk mencapai tujuan tersebut. karena sangat penting mencari ilmu demi masa depan yang cerah. pengawasan siswa oleh wali siswa membuat hasil prestasi anak tidak kunjung membaik. serta penilaian siswa oleh guru masih belum terkelola dengan baik karena masih berupa laporan catatan. metode sistem yang digunakan adalah model prototype. dengan observasi dan wawancara langsung dengan bagian siswa mengenai sistem penilaian di sekolah tempat penulis melakukan penelitian. hasil dari penelitian ini adalah aplikasi yang dapat dioperasikan pada smartphone android. aplikasi ini dapat memberikan informasi yang cepat dan terupdate secara otomatis dalam memperoleh informasi hasil belajar siswa. kata kunci: penilaian siswa, e-report card, metode prototipe, aplikasi. i. introduction education is one aspect that cannot be ruled out in life, the contribution of education to date is still expected to be improved, because this field can elevate the dignity of the nation and state, namely by producing human resources who can respond to world challenges. therefore, education will continue to be the government's main focus in realizing the intellectual life of the nation and [1]. the demands in today's millennial era require us to be able to keep up with increasingly rapid technological developments, especially with the presence of operating systems. android on smartphones is expected to be able to provide alternative solutions to solve the problems at hand. as the authority to supervise the learning process and results, the teacher provides reporting to students in the form of exams and student learning outcomes to the student's guardians to be evaluated by each student's guardian [2]. so that the child's learning outcomes and behavior can be monitored intensively by the student's guardian. one of the problems that arise from the case above is that there are still many parents who do not pay attention to their children's learning so that sometimes the learning process of their children is not closely supervised which results in a lack of motivation for children in learning because they feel less attention by their parents [3]. one of the reasons for the lack of parental attention to their children in learning. based on the problems stated above, the authors propose a medium and a solution to the problems that arise above, namely by making the application "e-reporting learning learners based on android (case study sd negeri cimahi ii)." ii. methods the system development method used in the preparation of this final project is to use the prototype method [4]. this method consists of several stages. the following are the stages of the prototype method in figure 1: 18 | vol.2 no.1, january 2021 figure 1 prototype stages 1. collection of needs at this stage, it is the stage that is carried out to collect data from various sources, namely by coming to the dinas to seek information about mutations [5]. then conduct interviews with transfer officers in the dinas. 2. build a prototype at this stage, it is a prototype application that will be made. 3. prototype evaluation at this stage, it is carried out by the teacher which evaluates the results of making sketches about student assessments. 4. system coding at this stage, the prototype is made using the php programming language with the concept of a code igniter framework [8]. 5. system testing after the system has become software, it must be tested before use. in testing this system, the black box testing method is used, in which the testing method is carried out on a program display that can run properly as desired, and white box testing which focuses on coding testing [9]. iii. results and discussion the process carried out at cimahi ii elementary school is not completely computerized and still uses bookkeeping [10]. therefore, the authors propose that the student assessment process can switch through the system. what is proposed by the author is based on android for users and based on websites for admins and teachers. the proposal aims to facilitate the transfer process that will be carried out at the dinas. below is a system design made by the author using astah [11]: a. use case diagram use case diagram describes some external actors and their relationship to the use case provided by the system [12]. the following is the use case diagram design: 1. usecase diagram below is a picture of a use case diagram for students. what is shown in figure 2: figure 2 usecase admin login diagram a. activity diagram activity diagrams describe a series of flow from activities, used to describe activities that are formed in an operation so that it can also be used for other activities such as use cases or interactions. the following is the activity diagram design: 1. admin login activity diagram admin login activity diagram is a description of the admin actor in accessing the system, where the admin fills in the username and password into the login menu, then the system will validate. 19 | vol.2 no.1, january 2021 figure 3 admin login activity diagram b. sequence diagram sequence diagram describes dynamic collaboration between some objects. its use is to show a series of messages sent between objects as well as interactions between objects. the following is the sequence diagram design: 1. sequence diagram perform admin login the sequence diagram for logging in is a picture of the interaction between menus, where the admin can log in. figure 4 admin login sequence diagram c. class diagram class diagram describes the static class structure in the system. the class represents something that is handled by the system. the following is the class diagram design shown in figure 5. figure 5 class diagram of student assessment d. system implementation system implementation is an explanation of how a program that has been made is run into a piece of hardware [13]. in the implementation process, software, namely the chrome browser, native, and the xampp application are used as virtual servers with apache and mysql server services installed [14]. the hardware used is a laptop with an intel core i3 processor with 4gb of ram. 1. implementation of the admin interface in this implementation, it displays the display of programs that have been run using a browser. the following are the results of the interface implementation: figure 6 admin login page figure 6 above is the admin login display image. where that page is the admin's first page to be able to enter the system by entering a username and password. 20 | vol.2 no.1, january 2021 figure 7 admin main page figure 7 above is the admin main page display. displays a list of menus for teachers, a list of students figure 8 teacher main page figure 8 above is the main display of the teacher's web which displays menus for inputting student scores. figure 9 display of student values figure 9 above is a display of the student scores input by the teacher. iv. conclusion in this final section, the writer will describe some conclusions that can be drawn and suggestions based on the findings of the research. in general, the authors conclude that: 1. with this application, teachers can enter student grades periodically through the e-reporting system for student assessments that can help teachers for assessment and archiving. 2. this student assessment e-reporting application is connected directly to a smartphone which is accessed by the student's guardian and can see the child's learning progress. 3. creating applications that can manage student grades and supervise student learning outcomes using the programming language php, javascript. reference [1] arman, a, (2017). sistem informasi pengolahan data penduduk nagari tanjung lolo , kecamatan tanjung gadang , kabupaten sijunjung berbasis web. jurnal edik informatika. vol. 2 (2) 163-170. [2] al fatta, hanif. 2007. analisis & perancangan sistem informasi untuk keunggulan bersaing perusahaan & organisasi modern. cv. andi offset. yogyakarta. [3] alexander f. k. sibero, 2011, kitab suci web programing, mediakom, yogyakarta. [4] hermawan, k., iskandar, a. a. and hartono, r. n. (2011) ‘development of ecg signal interpretation software on android 2.2’, in proceedings international conference on instrumentation, communication, information technology and biomedical engineering 2011, icici-bme 2011. doi: 10.1109/icici-bme.2011.6108621. [5] hilabi, s. s. and . p. (2018) ‘analisis kepuasan pengguna terhadap layanan aplikasi media sosial whatsapp mobile online’, buana ilmu. [6] imron. (2012:152). manajemen peserta didik berbasis sekolah. jakarta: bumi aksara. [7] kadir a. 2013. pemrograman database mysql. yogyakarta: mediakom [8] komputa (2015). pembangunan aplikasi child tracker berbasis assisted – global positioning system ( a-gps ) dengan platform android. jurnal ilmiah komputer dan informatika . [9] iswandy, eko (2015). sistem penunjang kebutuhan untuk menentukan penerimaan mahasiswa dan pelajar kurang mampu. jurnal teknoif, vol.3 (2). 2338-2724. [10] fujianto, a., & waspada, i. (2016). secara terpusat ( studi kasus cv . surya putra perkasa ). 9–10. [11] b. priatna, s. shofia hilabi, n. heryana, and a. solehudin, “aplikasi pengenalan tarian dan lagu tradisional indonesia berbasis multimedia,” sistematic, 2019, doi: 10.35706/sys.v1i2.1978. [12] tukino, shofia hilabi, s. and romadhon, h. (2020) ‘production raw material inventory control information system at pt. siix ems indonesia’, buana information technology and computer sciences (bit and cs), 1(1), pp. 8–11. doi: 10.36805/bit-cs.v1i1.681. [13] nugroho a. 2011. perancangan dan implementasi sistem basis data. yogyakarta: andi offset [14] baenil huda and saepul apriyanto, “aplikasi sistem informasi lowongan pekerjaan berbasis android dan web monitoring (penelitian dilakukan di kab. karawang),”buana ilmu, 2019, doi:10.36805/biv41.8.08. 21 | vol.2 no.1, january 2021 [15] widodo (2011:6). pemodelan sitem beroriwntasi objek dengan uml. graha ilmu, yogyakarta. paper title (use style: paper title) p-issn : 2715-2448 | e-issn : 2715-7199 vol.1 no.2, 10 july 2020 buana information tchnology and computer sciences (bit and cs) 33 | vol.1 no.2, 10 july 2020 buana information tchnology and computer sciences (bit and cs) password data authentication using a combination of md5 and playfair cipher matrix 13x13 bayu priyatna 1 technology information faculty of engineering computer universitas buana perjuangan karawang bayu.priyatna@ubpkarawang.ac.id ‹β› april lia hananto 2 school of computing, faculty of engineering universiti teknologi malaysia hananto1983@graduate.utm.my abstract—data security and confidentiality are the most important things that must be considered in information systems. to protect the cryptographic algorithm reliability it uses. md5 is one technique that is widely used in password data security issues, which algorithm has many advantages including. md5 has a one-way hash function so that the message has been converted to a message digest, and it is complicated to restore it to the original message (plaintext). in addition to the advantages of md5 also has a variety of shortcomings including; very easy to solve because md5 has a fixed encryption result, using the md5 modifier generator will be easily guessed, and md5 is not proper because it is vulnerable to collision attacks. the research method used in this study uses computer science engineering by conducting experiments combining two cryptographic arrangements. the results obtained from this study after being tested with avalanche effect technique get ciphertext randomness results of 44.12%, which tends to be very strong to be implemented in password data authentication. keywords— information systems, md5, playfire cipher. abstrak—keamanan dan kerahasiaan data merupakan hal terpenting yang harus diperhatikan pada sistem informasi. untuk menjahganya dibutuhkan kehandalan algoritma kriptografi yang digunakannya. md5 merupakan salah satu teknik yang banyak digunakan dalam masalah keamanan data password, yangmana algoritma ini memiliki banyak kelebihan diantaranya. md5 memiliki fungsi hash satu arah sehingga pesan yang telah diubah menjadi message digest (pesan ringkas), dan sangat sulit untuk mengembalikannya ke-pesan semula (plaintext). selain kelebihan md5 juga memiliki berbagai macam kekurangan diantaranya; sangat mudah di pecahkan karena md5 memiliki hasil enkripsi yang tetap, dengan menggunakan generator pengubah md5 akan dengan mudah ditebak, dan md5 kurang bagus karena rentan terhadap serangan collision attack. metode penelitian yang digunakan pada penelitian ini menggunakan rekayasa computer science dengan melakukan eksperimen penggabungan dua metode kriptografi. hasil yang didapat dari penelitian ini setelah di uji dengan teknik avalanche effect mendapatkan hasil keacakan cipherteks 44,79% yang cenderung sangat kuat untuk di implementasikan pada autentikasi data password. kata kunci—sistem informasi, md5, playfire cipher. i. introduction the issue of data security and confidentiality is one of the essential things of information systems [1]. the technique that can be used to maintain data content is to use cryptographic techniques [2]. cryptographic techniques aim to provide security services, including security, to manage data authentication such as passwords [3]. md5 is a one-way hash function designed by ron rivest with a 128-bit hash value. it says the one-way hash function because messages that have been converted to digest messages (full messages) are complicated to restore to the original message (plaintext) [4]. md5 is one of the one direction hash functions that is widely used to resolve the integrity of a file [5]. md5 is implemented in networks that produce 640-bit message digest [6]. md5 will be a highsecurity network for transferring data in cellular systems [7]. md5 is the result of encryption that is made easy to guess, although one-way md5 will be much easier to hack just using an md5 modifier generator will produce a straightforward match. the md5 encryption method is not suitable because it is vulnerable to collision attacks [8]. from these questions, this study discusses the authentication agreement on a password by combining the md5 method hash function and the application of playfire cryptography. ii. method the research method used in this research is engineering, namely theoretical computer science, where researchers use a cryptographic technique with the md5 method and combine it with the playfair algorithm using a 13x13 matrix. the systematic description of the process flow of this study is outlined in figure ii-1 as follows: 34 | vol.1 no.2, 10 july 2020 buana information tchnology and computer sciences (bit and cs) data password (plaintext) encryption 1 ciphertext 1 md5 encryption 2 playfire 13x13 key 1 ciphertext 2 encryption data password (ciphertext 2) plaintext 1 playfire 13x13 key 1 md5 plaintext 2 decryption decrypt 1 decrypt 2 fig. 1. research methods a. md5 algorithm the hash function is widely used in md5 and sha cryptography. in this article, the hash function that is used by md5 algorithm [9]. md5 accepts input in the form of messages of any size and produces a message digest that is 128 bits in length [10]. md5 processes 512-bit blocks, divided into 16 32-bit sub-blocks. the algorithm output is set to 4 blocks, each measuring 32 bits, which, after being combined, will form a 128-bit hash value [11]. the steps in making a message digest, in general, are as follows: 1. adding padding bits. 2. add the original message length value. 3. initialize the md5 buffer. 4. processing messages in blocks of 512 bits. b. playfair matrix 13x13 formation of a 13 x 13 playfair matrix table of keys that have been entered, on the formation of a key consisting of letters, numbers and symbols, for example, the example key "akum@ululu5". the first step is a key that consists of numbers, letters or symbols should not have more than one appearance if there are these things, then eliminate numbers, letters or symbols that have similarities. so the key from "akum@ululu5" becomes "akum@ul5"[2];[12]. in table iii-2 is a matrix formed by the key "akum@u15": fig. 2. playfair matrix 13x13 [2];[13]. c. playfire encryption algorithm before doing the encryption process, the plaintext to be encrypted is set as follows: 1. all characters and spaces not included in the alphabet must be removed from the plaintext (if any). 2. if there is a letter j in the plaintext make changes with the letter i. 3. the plaintext, which is the original message, is arranged according to the letter pair (bigram). 4. when there are the same pair of letters, do it change one of the letters of the letter pairs with the letter z or x insert using the letter x because the letter x is very minimal at all in bigram, unlike the letter z, for example, is the word fuzzy. 5. if the letters in the plaintext have an odd number, then select additional messages then add at the end of the plaintext. other notes can be chosen, for example, the letter z or x [14]. d. playfire decryption algorithm following are the stages of the playfair cipher algorithm: 1. if there are two letters on the same key line, then each letter is changed using the message to the left. 2. if there are two letters in the same column, later each letter is changed by the message above. 3. if two letters are not located in the same row and column, then replace them with the word in the intersection of the first row of words with the column letter two. 4. later the second letter is changed using the word at the vertex of the rectangle made from the letter used [15]. iii. result and discussion a. interface system the application interface built can be seen in figure 3: 35 | vol.1 no.2, 10 july 2020 buana information tchnology and computer sciences (bit and cs) fig. 3. form login admin b. database system on this system, a database is created with one of the table names db_user as the sample used. here is figure 4 that shows the attributes of the user table and figure 5 shows the results of encryption: fig. 4. user table structure fig. 5. password encryption results c. ciperteks randomness test in the ciphertext randomness test, 30 trials were performed with random sample parameters based on the same data size, character length, and key. the results of the experiment using the calculation of the avalanche effect method with the formula: in general, bits in ciphertext will change from the number of bits in plaintext by 50%. the avalanche effect is approved well if the resulting bit change gets 45-60% (about half): the more changes that occur, the more difficult cryptographic algorithms to be completed or have high complexity. landslide implementation effect of the number of bit changes obtained from the xor calculation from the plaintext and ciphertext distribution to binary numbers, then prove the combined md5 with the playfair 13x13 algorithm. graph of landslide effects: table iii-1 comparison of ciphertext randomness test results no data password plaintext length (bit) afvalanche effect 1 sample 1 142 43,18 2 sample 2 482 41,31 3 sample 3 993 45,07 4 sample 4 1287 47,06 5 sample 5 1.981 47,51 6 sample 6 2.212 40,87 7 sample 7 3.580 45,34 8 sample 8 3.780 47,31 9 sample 9 4.324 44,96 10 sample 10 5.520 45,15 11 sample 11 5.804 46,93 12 sample 12 6.916 48,77 13 sample 13 7.882 44,92 14 sample 14 8.576 45,01 15 sample 15 18.796 45,97 16 sample 16 10.940 38,81 17 sample 17 16.980 43,62 18 sample 18 12.176 45,43 19 sample 19 31.992 44,94 20 sample 20 19.840 39,61 21 sample 21 33.372 40,14 22 sample 22 35.796 39,96 23 sample 23 25.692 44,40 24 sample 24 38.680 44,80 25 sample 25 58.286 44,48 26 sample 26 70.212 40,10 27 sample 27 95.890 44,73 28 sample 28 121.748 44,49 29 sample 29 138.780 44,15 30 sample 30 141.468 44,63 avalanche effect average score 44,12 fig. 6. avalanche effect test results conclusion based on the results of research conducted, password authentication data security can answer the hypothesis at the beginning of the study by applying the md5 algorithm and then combining it with playfair 13x13, which can improve 36 | vol.1 no.2, 10 july 2020 buana information tchnology and computer sciences (bit and cs) the security data on the previous md5 password algorithm, with the avalanche effect test results of 44.79%. besides having the strength of the md5 encryption technique with a combination of playfire 13x13 can cover the deficiencies found in the md5 method. references [1] inayatullah, “analisis penerapan algoritma md5 untuk pengamanan password,” j. algoritm., vol. 3, no. 3, pp. 1–5, 2017. [2] d. kurniawan and b. priyatna, “pengamanan data berbasis mobile android dengan penggabungan linear feedback shift register (lfsr) dan modifikasi matriks kunci algoritma kriptografi playfair cipher,” j. telemat. mkom vol, vol. 10, no. 1, 2018. [3] b. priyatna and a. l. hananto, “zfone security analysis of video call service using general nework design process method ( gndp ).” [4] a. l. hananto, a. solehudin, a. s. y. irawan, and b. priyatna, “analyzing the kasiski method against vigenere cipher,” arxiv prepr. arxiv1912.04519, 2019. [5] n. hayati, m. a. budiman, and a. sharif, “implementasi algoritma rc4a dan md5 untuk menjamin confidentiality dan integrity pada file teks,” sinkron, vol. 1, no. 2, 2017. [6] d. sharma, p. sarao, and s. dudi, “implementation of md5-640 bits algorithm,” int. j. adv. res. comput. sci. manag. stud., vol. 3, no. 5, pp. 286– 293, 2015. [7] a. qashlim and rusdianto, “implementasi algoritma md5 untuk keamanan dokumen,” j. ilm. ilmu komput., vol. 2, no. 2, pp. 10–16, 2016. [8] t. f. prasetyo and a. hikmawan, “analisis perbandingan dan implementasi sistem keamanan data menggunakan metode enkripsi rc4 sha dan md5,” infotech j., vol. 2, no. 1, 2016. [9] m.-j. wang and y.-z. li, “hash function with variable output length,” 2015 int. conf. netw. inf. syst. comput., pp. 190–193, 2015. [10] z. musliyana, t. y. arif, and r. munadi, “peningkatan sistem keamanan autentikasi single sign on (sso) menggunakan algoritma aes dan one-time password studi kasus: sso universitas ubudiyah indonesia,” j. rekayasa elektr., vol. 12, no. 1, p. 21, 2016. [11] i. hmac-sha and h.-p. ipsec, “implementasi hmacsha1,tripledes-cbc,hmac-md5-96 pada ipsec,” no. may, 2016. [12] d. kurniawan, a. l. hananto, and b. priyatna, “modification application of key metrics 13x13 cryptographic algorithm playfair cipher and combination with linear feedback shift register (lfsr) on data security based on mobile android,” int. j. comput. tech. -–, vol. 5, no. 1, pp. 65–70, 2018. [13] a. l. hananto and a. r. priyatna, bayu, “android data security using cryptographic algorithm combinations,” int. j. psychosoc. rehabil., vol. volume 24, no. issue 7, pp. 3307–3318, 2020. [14] r. m. marzan and a. m. sison, “an enhanced key security of playfair cipher algorithm,” in proceedings of the 2019 8th international conference on software and computer applications, 2019, pp. 457–461. [15] e. h. nurkifli, “modifikasi algoritma playfair dan menggabungkan dengan linear feedback shift register ( lfsr ),” vol. 2014, no. sentika, 2014. paper title (use style: paper title) issn : vol.1 no.1, 01 january 2020 buana information tchnology and computer sciences (bit and cs) 19 | vol.1 no.1, 01 january 2020 buana information tchnology and computer sciences (bit and cs) implementation of the futsal field ordering platform using the ucd method agustia hananto 1 technology information master of computer science university of budi luhur agustia.hananto@gmail.com muhamad mammun 2 information systems study program school of engineering and computer science universitas buana perjuangan karawang mamun20@gmail.com ‹β› nurhayati 3 technology information master of computer science university of budi luhur hayatinur10@gmail.ac.id abstract—the development of information technology is exploding. we cannot separate the need for information from the use and use of computers. with a computerized information system, the work done will be more effective and accurate. karawang futsal is a sports venue in the karawang regency. using the futsal ordering system is still manual, the data input system which is still recording in the ledger, making reports is not accurate because of frequent miscalculations that result in making reports not on time because all processes are done. therefore, with the existence of a computer system, all the needs for everything in the karawang regency futsal will run. keywords—information systems, computerized, futsal. i. introduction information technology is a technology that is developing. the information available can take place, and. information that is accurate and can be accessed by anyone, anywhere and anytime, with information systems using computers as a medium that makes it easier for someone to manage data. good data and information management are very important for the needs of an organization, those related to business. one example is the futsal field ordering system. futsal court booking is a booming business that provides futsal field booking services. business processes in place for futsal field bookings still require customers to come in making an order and arrange the desired booking schedule. so customers do not know the empty schedule. every day the clerk records the orders from the customers in the order book. on the day of the order, the customer makes an order for payment. this can cause errors in recording. this manual field booking system is inconvenient for the field user and becomes less efficient in terms of time, energy, and cost because the user must go to any existing futsal venue to check the schedule and field booking. based on these constraints, it is very much needed a system with webbased futsal field booking design, in terms of accurate validation for scheduling and field booking problems. web applications are easier to access. a website can be accessed from anywhere as long as there is an internet network. this application helps consumers to see the field schedule and can order according to the desired time. this application is also designed so that owners of futsal venues can manage and manage their field schedules. using this system is able to manage futsal field bookings,, and. this web-based futsal field booking application is expected to help users to provide information about the field and make reservations and. ii. theoretical basis a. system a system is a unit comprising components or elements that are linked and arranged in such a way as to facilitate the flow of information that serves to achieve a goal [5], the system is a collection of sub-systems, elements, procedures that are integrated to achieve certain goals, such as target information or goals. meanwhile, the system is a collection of components that work together to achieve a goal [6]. b. information information is very important for the company in making every decision, information comes from the ancient french language, information in 1387 which was taken from the latin information which means "outline, concept, idea" [3]. information is data processed into a form that is more useful and more meaningful for those who receive it, while the data is a source of information that describes a real event information is data that has been organized and has had uses and benefits [4]. c. black-box testing the tester uses behavioral tests (also called black-box tests), often used to find bugs in high-level operations, at feature levels, operational profiles, and customer scenarios. the tester can make functional black box testing based on what the system has to do. behavioral testing involves a detailed understanding of the application domain, the business problems that are solved by the system and the system's mission. behavioral tests are best by testers who understand system design, at least at a high level so they can find common bugs for this design [2]. d. ucd (user centered design) ucd is the user's relation to the whole process. users/users not only provide input on the design concept but also must be involved in all aspects, including the stages of implementation in the system that will affect their activities. the user is also involved in the initial testing and evaluation and design. but depending on the complexity of the system to be built, there is [1]. iii. research methods before the research method used is the waterfall model. the waterfall comprises several stages of activity flow that goes one direction from the beginning to the end of the 20 | vol.1 no.1, 01 january 2020 buana information tchnology and computer sciences (bit and cs) system development project. the waterfall sdlc model provides a sequential or sequential software life-flow approach starting from the analysis, design, coding, testing, and support stages. here is a picture of the waterfall model: pelanggan menentukan lapangan futsal memesan lapangan futsal melakukan pembayaran menyediakan lapangan melakukan konfirmasi pembayaran admin fig. 1. use case diagram a. analysis the author analyzes the needs of the application to be built. the analysis is done by direct observation of research objects. it made observations to retrieve data approved by the object of research for study by researchers. b. the design the design is used to make an initial description of the application to be made, in system modeling, the application interface for photocell field ordering in karawang. c. making program code at this stage, the authors do the application coding by the design that has been made. d. testing they carry the test out to test both the logic and functionality of the application that has been made, to ensure the application runs as desired and according to its function. in this study, the study conducted testing with the black box method. e. maintenance i do maintenance after the application is running by checking the running of the application and backing up the data in a scaled way. . iv. result and discussion a. analysis of current system procedures explain the flow of new systems running in the form of information flow patterns that occur through documents, reports, processes or procedures that occur in the new system that is running. b. system architecture i built this system to provide information about futsal field reservations in the city of karawang through a media website. this futsal field data object is managed by an admin: fig. 2. system architecture c. database 1. class diagram class diagram is the most important element in objectoriented systems, the class describes a building block system. class diagram has other features and characteristics, while those listed in this system are those that are related to the design of a futsal field reservation system, the following class diagram on the futsal field order information system: login +username +password +input_data() member +username +password +input_data() +alamat +nama +hp pemesanan +kode_pemesanan +tgl_pemesanan +input_data() +status +member lapangan +kode_lapangan +nama_futsal +input_data() +jenis_lantai +nama_lapangan +luas_lapangan +deskripsi_lapangan +harga +simpan() +batal() +simpan() +simpan() konfirmasi pembayaran +kode_pemesanan +bank_tujuan +input_data() +no_rekening_pengirim +bank_pengirim +atas_nama_rekening +jumlah +simpan() 1 1 1 1 1m m m fig. 3. class diagram d. user interface 1. website mann page this main output page is the first page that will appear when the user enters the website address of the karawang regency futsal field booking platform website. this main page comprises several main menus, the home menu, today's booking menu, the schedule menu, the booking menu and the login menu, which are enabled to make it easier for users to find out futsal field booking information. next website home page display: 21 | vol.1 no.1, 01 january 2020 buana information tchnology and computer sciences (bit and cs) fig. 4. admin page 2. admin login page admin login input page to use all the admin features must first log in, next is the admin login page display: fig. 5. login page 3. member registration page on the input page of this member registration, a data form will appear that the user must fill in if he wants to register as a member. the appearance of member registration pages is: fig. 6. registration page 4. add booking page on the addition data booking input page, a data form will appear that the user must fill in if he wants to book a futsal field. the appearance of the additional data booking page is: fig. 7. add booking page e. testing testing this web-based futsal field ordering information system software using the black box method. black box testing focuses on the functional requirements of the admin login form and member login form software. table 1. testing web grade test item test test level testing types login login admin integrasi black box login user integrasi black box testing pengisian registrasi user integrasi black box data testing pengisian transaksi integrasi black box conclusion from the results of field research and the website creation process, the authors conclude: 1. this photocell field ordering system can already be used. implementing this web-based information system can make it easier for customers and managers of futsal fields to get information relating to the futsal venue, futsal field reservations. 2. in this futsal field booking website information systems can print proof of payment for futsal field bookings to reduce fraud in payments.unnumbered footnote on the first page. references [1] bayu priyatna, “penerapan metode user centered design (ucd) pada sistem pemesanan menu kuliner nusantara berbasis mobile android,” account. inf. syst. penerapan, pp. 17–30, 2018. [2] black, m. j. & hawks, h. j., 2009.medical surgical nursing: clinical management for continuity of care, 8th ed. philadephia: w.b. saunders company [3] mulyanto agus. 2009. sistem informasi konsep dan aplikasi. pustaka pelajar. yogyakarta. [4] suryana, taryana dan koesheryatin. 2014. aplikasi internet menggunakan html, css, & javascript.jakarta: pt elex media komputindo. [5] wahyono, & teguh. 2004. sistem informasi konsep dasar, analisis, desain dan implementasi. 22 | vol.1 no.1, 01 january 2020 buana information tchnology and computer sciences (bit and cs) graha ilmu, yogyakarta. [6] suryana, taryana dan koesheryatin. 2014. aplikasi internet menggunakan html, css, & javascript.jakarta: pt elex media komputindo. paper title (use style: paper title) issn : vol.1 no.1, 01 january 2020 buana information tchnology and computer sciences (bit and cs) 8 | vol.1 no.1, 01 january 2020 buana information tchnology and computer sciences (bit and cs) production raw material inventory control information system at pt. siix ems indonesia tukino 1 information systems study program school of engineering and computer science universitas buana perjuangan karawang tukino@ubpkarawang.ac.id shofa shofia hilabi 2 information systems study program school of engineering and computer science universitas buana perjuangan karawang shofa.hilabi@ubpkarawang.ac.id ‹β› heri romadhon 3 information systems study program school of engineering and computer science universitas buana perjuangan karawang si15.heri.romadhon@ubpkarawang.ac.id abstract—application of web-based raw material inventory control information system with a bill of materials (bom) using the php and mysql programming language as a database, and using the sdlc livestock device engineering method with stages of planning, design, implementation, and testing. i create a recording application that provides information about the availability of raw material reducing the error in calculating the amount of raw material based on the bill of material. with this application, it can help to purchase in determining the number of raw materials needed for production based on the master bill of material data. keywords— inventory, bill of material, information system. i. introduction this the production process is the core activity of a manufacturing company. in the production process, they require a company to produce a quality product following consumer desires. to conduct production activities, good raw materials must be available and following the company's production needs. therefore, determining the raw material inventory and is a very important activity in a production process. planning for the right raw material inventory is very supportive in the smooth production process. the smooth production process is very important for the company because it is very influential on the level of sales and profits got by the company. factors that influence the smooth production process are the availability of raw materials that will be processed in the production process. if raw material inventory is not available with the required amount or raw material is late until the company, then it will have a bad influence on the company that is affecting the company's profits, this is because of costs incurred for the company running out of inventory which results in lost opportunities benefit because consumer demand cannot be served and the production process is interrupted. in the research the author will conduct that is research on controlling raw materials in electronic products in the siix ems (electronic manufacturing services) company. the main raw materials for the electronic products are pcb (printed circuit board), capacitors, resistors, leds (liquid crystal display) and diodes. description of raw material into finished goods presented by pt. sipx ems indonesia in the form of (bom) bill of materials is a special table that shows how raw material changes into finished products. in the bill of material will show each product has a composition of any raw material and the size or amount needed to make one finished product. with the problems that occur in controlling the residual and the need for the use of raw material production, the author intends to analyze the system design information system inventory control material production row in p.t. siix ems indonesia. ii. theoretical basis a. information systems in the research process information system design requires an understanding of information systems, there are several opinions according to experts as follows: information systems are systems that can be defined by collecting, processing, storing, analyzing, distributing, information for specific purposes. like other systems, an information system consists of inputs (data, instructions) and outputs (reports, calculations) " [3]. information systems is a system within an organization that meets the needs of daily transactions that support the organization's operational functions that are managerial with the strategic activities of an organization to be able to provide certain external parties with the reports needed "[4]. in simple information systems of understanding above have in common that the information system will produce a report, meaning a process that produces a report or data. b. system analysis in a study a researcher must find the existing problems and solve those problems with certain methods of the system running and aims to find out all the problems that occur and facilitate in carrying out the next stage through analysis that has been made, according to experts the analysis of the system can be defined as the following: argues that system analysis is a process of reviewing an information system and dividing it into its constituent components for later research so that known problems and needs that will arise so that it can be reported in full and proposed improvements to the system " [5]. analysis of the system is a method for finding solutions to existing system problems by grouping existing 9 | vol.1 no.1, 01 january 2020 buana information tchnology and computer sciences (bit and cs) components into smaller components so that the solutions found by system requirements [6]. from the above definition, quotations have in common in the understanding that analysis is the process of solving the problem means that the existing problems will be solved by a particular method to get a solution to the problem. c. waterfall someone who conducts research certainly has a method of solving problems in a study, one of which is the waterfall method of various methods that will also be used by the author. this method is a classic method but is still often used in research methods, experts are precise about this method as follows: waterfall is one method of developing information systems that are systematic and sequential, meaning that each stage in this method is carried out sequentially and continuously [7]. the waterfall model as one of the basic theories and as required to be studied in the context of the software life cycle, is a life cycle that consists of starting the life phase of the software before it occurs until post-production [8]. the description of the waterfall method defined by the experts above can be drawn in outline by the authors of the method which is systematically structured from the phase before it starts to the phase after the information system starts. the stages of the waterfall method can be seen in the figure below according to [9]. fig. 1. waterfall model [9]. d. raw material inventory in a company, every operational manager is required to be able to manage and hold inventory to create effectiveness and efficiency. inventory of raw materials is inventory of raw materials has an important position in the company because the supply of raw materials is a very large influence on the smooth production process [13]. inventories of raw materials are inventories are stored materials or goods that will be used to fulfill certain purposes, for example, for use in the production or assembly process, for resale, or parts of an equipment or machine". from the above understanding, it can be understood that raw materials are raw materials, semi-finished materials, and materials that will be processed in the production process to become a finished product [14]. iii. research methods in this study, there is further research that contains the stages. this framework is the steps or stages that will be carried out by researchers in the settlement that will be discussed. is a flow cart picture. fig. 2. research framework iv. result and discussion a. use case system diagram to illustrate the system activities that will be made, the author uses modeling with use case diagrams. the use case diagram is used to find out what functions existing in the system and anyone who may use these functions. the use case diagram for the following raw material inventory control systems: fig. 3. use case diagram of raw material inventory control system 10 | vol.1 no.1, 01 january 2020 buana information tchnology and computer sciences (bit and cs) b. class diagram after the use case diagram, the writer makes a class diagram aimed at a container that describes the structure of objects in the system formed from the relationships between classes. the class diagram is: fig. 4. class diagram c. implementation 1. login page the login page is the start page that appears when opening a user page (backend). the function of the page is for the user to enter the main web page application: fig. 5. login page 1. main page fig. 6. main page conclusion based on the analysis that has been done, the authors can conclude that consumer satisfaction is a major factor in winning increasingly fierce industry competition. consumer satisfaction can be achieved in several ways including quality products, competitive prices, and timely delivery. therefore, a company needs to pay attention to the control and planning of raw material inventory (raw material) to achieve the effectiveness and efficiency of data processing to maintain smooth production and increase customer satisfaction. to be able to optimize the inventory function, companies must make plans in the procurement of raw materials. the planning must be following the production needs for each boast with the existence of a raw material control system that has been made, is expected to facilitate the admin in conducting raw material relations, search for data to make reports, so that the creation of better jobs. as for the discussion of the previous chapters conclusions can be drawn as follows: 1. the raw material control system at pt. six ems indonesia, especially for calculating raw material production needs, still use excel files and different storage databases. this causes the operational processes that are not yet maximized. 2. calculation of raw material requirements in the application using the bill of material multiplication. 3. in the application, a form is provided for request input or forecasts from the customer and then automatically calculates the raw material production requirements. references [1] swastika, i.p.a. & putra, i.g.l.a.r. 2016. audit sistem informasi dan tata kelola teknologi informasi. yogyakarta: cv. andi offset. [2] ahmad, rohani. 2010. pengelolaan pembelajaran. jakarta: pt rineka cipta. [3] sutarman. 2012. buku pengantar teknologi informasi. jakarta: bumi aksara. [4] tata sutabri. 2012. konsep sistem informasi. andi: yogyakarta. [5] wahana, komputer. 2010. shourtcourse sql server 2008 express. yogyakarta: andi offset. [6] whitten & jeffery , bentley. 2009. system analysis and design methods. new york: the mcgraw-hill companies, inc. 11 | vol.1 no.1, 01 january 2020 buana information tchnology and computer sciences (bit and cs) [7] nasution, ruslan efendi. 2012. implementasi sms gateway in the development web based information system schedule seminar tesis. lampung : unila. [8] rizky, soetam. 2011. konsep dasar rekayasa perangkat lunak. jakarta: prestasi pustaka. [9] roger, s & pressman, ph.d. 2012. rekayasa perangkat lunak .yogyakarta: andi. [10] budiman, agustiar. 2012. pengujian perangkat lunak denganmetode black box pada proses pra registrasi uservia website. makalah hal:4. [11] al-bahra bin ladjamudin. 2013. analisis dan desain sistem informasi. yogyakarta : graha ilmu. [12] shalahuddin, m. 2013. rekayasa perangkat lunak terstruktur dan berorientasi objek. bandung: bandung. [13] rangkuti f, (2007). manajemen persediaan “aplikasi di bidang bisinis”. jakarta : raja grafindo persada [14] eddy, herjanto. (2007. manajeman produksi dan operasi . jakarta :pt. gramedia wiarsarana indonesia. [15] baridwan, zaki. 2011. intermediate accounting edisi 8. yogyakarta : bpfe. [16] assauri, sofjan. 2008. majemen produksi dan operasi. jakarta paper title (use style: paper title) 22 | vol.2 no.1, january 2021 issn : 2715-2448 | e-issn : 2715-7199 vol.2 no.1 january 2021 buana information technology and computer sciences (bit and cs) application for submission of research recommendations and practice work web based shofa shofiah hilabi 1 information system, faculty of engineering and computer science universitas buana perjuangan karawang, indonesia shofa.hilabi@ubpkarawang.ac.id arip solehudin 2 program studyteknik informatikafakultas ilmu komputer, universitas singaperbangsa karawang arip.solehudin@staff.unsika.ac.id syahri susanto 3 school of oil technic balongan oil and gas academy indramayu, indonesia syahri28@gmail.com ‹β› abstract — application submission of recommendations for research and job training which is often late hampers students who will carry out research and practical work activities that will be carried out at institutions in karawang, making students have to spend time visiting the kesbangpol office where sometimes the completion of the letter is not clear when the completion completed, and also because of the head of the office's signature problem. the research method used is to use the waterfall method which in this method includes needs analysis, design, implementation, verification, and maintenance. therefore it is hoped that after the design and recommendation system of a recommendation letter system / tools proposed by the author, it is useful for efficient time for students who will submit a recommendation letter for research and practical work to agencies in karawang regency and students do not have to go directly to the kesbangpol office. karawang regency to submit a recommendation letter keywoard: applications, recommendations, research letters, practical work. abstrak — aplikasi pengajuan surat rekomendasi penelitian dan kerja praktek yang sering kali terlambat menghambat mahasiswa yang akan melakukan kegiatan penelitian dan kerja praktek yang akan dilakukan di instansi yang ada di karawang, membuat mahasiswa harus menghabiskan waktu untuk mendatangi kantor kesbangpol yang mana kadang penyelesaian surat tersebut tidak jelas penyelesaianya kapan selesai, dan juga karena kendala tandatangan kepala kantor. metode penelitian yang digunakan adalah menggunakan metode waterfall dimana dalam metode ini meliputi analisis kebutuhan, desain, implementasi, verifikasi, dan pemeliharaan. maka dari itu di harapkan setelah di rancangnya dan sistem rekomendasi sebuah sistem/tools surat rekomendasi yang di usulkan oleh penulis, bermanfaat mengefisiensi waktu mahasiswa yang akan mengajukan surat rekomendasi penelitian dan kerja praktek ke instansi yang di kabupaten karawang dan mahasiswa tidak harus mendatangi langsung kantor kesbangpol kabupaten karawang untuk mengajukan surat rekomendasu. kata kunci: aplikasi, rekomendasi, surat penelitian, kerja praktek. i. introduction local government has a function in serving the community in administrative and bureaucratic matters. local governments have various regional work units (skpd) that carry out their main tasks for the benefit of the community. regional apparatus work units (skpd) are elements of regional government administration that to achieve success need to be supported by excellent planning by the organization's vision and mission [1]. a letter of recommendation is bidding made by a certain leader or official that contains information about a person's situation based on authentic data available because the party concerned has asked for his interest [2]. one of the karawang regency regional work units, namely the office of national unity and politics (kesbangpol) of karawang regency has the main task of assisting the regent in carrying out regional government affairs based on the principle of autonomy, namely in the preparation and implementation of regional policies in the field of national unity and politics [3]. the national unity and political agency also have tasks including collecting data on the names of the management structure of the secretariat address for all community organizations, non-governmental organizations located in karawang regency and also providing guidance for community organizations, non-governmental organizations in karawang. 1 (one) year and kesbangpol always takes action against problems that occur between ormas and other parties using mediating for the sake of unity and integrity as well as conduciveness. the national unity and political agency also have a duty to facilitate or provide letters of recommendation/introduction to student activities or institutions that will conduct research or carry out practical work. the recommendation letter will be shown to government agencies in karawang, which if students or institutions that are going to make a recommendation letter must first come to the kesbangpol office to find out what there are only requirements needed to submit a recommendation letter. the process of making a recommendation letter from the kesbangpol office is relatively short and easy because there is already a certain format for employees. however, the recommendation letter must be signed by the head of kesbangpol karawang, which sometimes the head of kesbangpol has duties outside the office so that when asked to sign it is quite difficult to contact and difficult to ask mailto:shofa.hilabi@ubpkarawang.ac.id mailto:arip.solehudin@staff.unsika.ac.id 23 | vol.2 no.1, january 2021 for free time to sign recommendation letters that have been submitted by students or institutions. that is to save time, therefore researchers will create a web-based information system which will make it easier for students or institutions that will submit letters of recommendation addressed to government agencies located in karawang, students or institutions only need to open a website which will later be designed by the author, in submitting a letter of recommendation. if a student/institution is going to submit a recommendation letter, the student/institution must have the requirements needed by kesbang employees, including, ktp, kta, a letter from the university, if a web-based information system has been designed, the office of national unity and student politics and institutions is only sufficient. scanning and entering it into the system that has been designed by the author, then the student who will submit also lists which agency it is intended for, and from what date and until what month the student will conduct research / practical work activities, then if the letter has been approved by the employee a notification will appear in the system, and students just need to print the proposed recommendation letter for submission, to the place of the agency to be examined or to do practical work. therefore, the author will design a "application for submission of recommendations for submission of research and job training case studies at the office of national unity and politics of karawang regency" so that students or institutions that will make recommendation letters do not have to come to the office directly and it is easier and more efficient. time. ii. method the results of this study are using the waterfall method wherein this method includes needs analysis, design, implementation, verification, and maintenance, producing some of the data needed by the author to design a web-based recommendation letter submission system, all of the research results are obtained by observing. directly at the place of researchers and interviews with several users related to research and data collection to achieve this research [4]. the following is figure 1. research flow figure 1. research flow systems development methodology the methodology used in system development for the design and development of this system is the waterfall methodology [5]. the following is figure 2. the stages of the waterfall methodology figure 2.stages of the waterfall methodology several stages of the waterfall methodology, namely: 1. requirements analysis, collecting the complete needs then analyzed and defined the needs that must be met by the program to be built [6]. data collection was carried out by the author through observation, interviews and documentation. 2. design, in this stage the developer will produce an overall system and determine the software flow to a detailed algorithm [5]. 3. implementation is the stage where the entire design is converted into program code. the resulting program code is still in the form of modules that will be integrated into a complete system [7]. 4. integration & testing. this stage is carried out by combining modules that have been made and this testing is carried out to find out whether the software made is in accordance with the design and functions of the software [8]. 5. verification is the client or user tests whether the system is approved [9]. 6. operation and maintenance, namely the installation and repair process of the system as approved [10]. iii. results and discussion the application is designed using php and mysql database for data storage. before entering the coding stage, the design and flow of the proposed system will be made. the following are some of the stages in system and software design: a. use case diagram so to illustrate the system activity that will be designed by the author using modeling using use case diagrams, it is used to find out what functions are in it, and who uses these functions [11]. the following is a figure 3.use case is the system recommended by the author. 24 | vol.2 no.1, january 2021 figure 3. here is a use case diagram of a recommended application a. activity diagram activity diagrams are used to describe various activity flows in a system that is being designed and how each flow begins [12]. the following is a figure 4.a diagram of the proposed recommendation system activity. figure 4.activity diagram of the proposed recommendation system figure 5. class diagram of the recommendation system system requirements analysis system requirements analysis is used to identify new system requirements [14]. system requirements include user needs and admin needs and analysis of platform requirements. application for submitting letters of recommendation. 1. admin page a. login b. agency data c. university data d. official data e. transaction data f. report data 2. member page a. registration b. login c. recommended data application implementation the following is an implementation of the interface on the initial display before entering login [15]. the image below will display a dashboard displaying several menus that can be used by the admin. then the admin can use the menu as needed by the admin. b. class diagram class diagrams are used to explain the structure of the system in terms of defining the classes that will be made to build a system [13]. the following is figure 5. class diagram of the recommendation system. figure 6. here is the main admin page the image below will display a confirmation menu of the recommendation system process used by the admin. figure 7. here is the main admin page for system confirmation 25 | vol.2 no.1, january 2021 submission in figure 7. below below will display a dashboard displaying several menus that can be used by the user. then the user can use the menu as needed by the user. figure 8 this is the user's main page the image below displays a dashboard displaying a recommendation menu that can be used by the user in submitting a recommendation letter. figure 9. below is the user recommendation menu page iv. conclusion based on the description and overall discussion, the application for submitting research recommendations and practical case study work at the national unity and political office of karawang regency, conclusions can be drawn, as follows: 1. to optimize the submission of recommendation letters, this application is made to provide more effective information, and students who submit recommendation letters at kesbangpol can submit recommendation letters for more than one submission of research recommendation letters or practical work. 2. creating a web-based application of recommendation letters can make it easier for the admin to verify or validate the submission of recommendation letters so that students do not fill the kesbangpol office. daftar pustaka [1] c. nisak, p. fitri, and a. kurniawan, “sistem pengendalian intern dalam pencegahan fraud pada satuan kerja perangkat daerah (skpd) pada kabupaten bangkalan,” jaffa, 2013. [2] s. rachmatullah and a. p. wijaya, “rekomendasi disposisi surat dengan metode naïve bayes pada arsip surat di kantor bakorwil kabupaten pamekasan,” j. comput. inf. technol., 2019. [3] y. efyanti, “peran kesbangpol linmas dalam pembinaan organisasi sosial politik dan organisasi kemasyarakatan,” islam. j. ilmu-ilmu keislam., 2019, doi: 10.32939/islamika.v18i02.311. [4] b. huda, “sistem informasi data penduduk berbasis android dan web monitoring studi kasus pemerintah kota karawang (penelitian dilakukan di kab. karawang),” buana ilmu, 2018, doi: 10.36805/bi.v3i1.456. [5] m. s. rosa a.s, “model waterfall,” 2016. 2016. [6] d. c. p. b. stmik nusa mandiri jakarta and i. s. stmik nusa mandiri jakarta, “perancangan sistem informasi balai kesehatan tni al pangkalan jati menggunakan metode waterfall,” evolusi j. sains dan manaj., 2018, doi: 10.31294/evolusi.v6i1.3536. [7] r. susanto and a. d. andriana, “perbandingan model waterfall dan prototyping,” maj. ilm. unikom, 2016. [8] i. binanto, “analisa metode classic life cycle ( waterfall) untuk pengembangan perangkat lunak multimedia,” j. univ. sanata dharma yogyakarta, 2014, doi: 10.13140/2.1.1586.4968. [9] a. abdurrahman and s. masripah, “metode waterfall untuk sistem informasi penjualan,” inf. syst. educ. prof., 2017. [10] r. susanto and a. d. andriana, “perbandingan model waterfall dan prototyping untuk pengembangan sistem informasi,” maj. ilm. unikom, 2016, doi: 10.34010/miu.v14i1.174. [11] b. priyatna, s. shofia hilabi, n. heryana, and a. solehudin, “aplikasi pengenalan tarian dan lagu tradisional indonesia berbasis multimedia,” systematics, 2019, doi: 10.35706/sys.v1i2.1978. [12] a. l. hananto and a. r. priyatna, bayu, “android data security using cryptographic algorithm combinations,” int. j. psychosoc. rehabil., 2020. [13] s. aripiyanto, “pengembangan prototipe sistem informasi monitoring hardware it berbasis web dengan metode kano dan model view controller : studi kasus pada pt. kalbe morinaga indonesia,” techno xplore j. ilmu komput. dan teknol. inf., 2018, doi: 10.36805/technoxplore.v2i2.303. [14] f. nugraha, “analisa dan perancangan sistem informasi perpustakaan,” j. teknol. inf. pendidik. itp, 2014. [15] c. a. pamungkas, “pengantar dan implementasi basis data,” in pengantar dan implementasi basis data, 2017. paper title (use style: paper title) p-issn : 2715-2448 | e-issn : 2715-7199 vol.1 no.2, 10 july 2020 buana information tchnology and computer sciences (bit and cs) 42 | vol.1 no.2, 10 july 2020 buana information tchnology and computer sciences (bit and cs) combination of hill cipher algorithm and caesar cipher algorithm for exam data security agung susilo yuda irawan 1 study program technical information faculty of computer science, universitas singaperbangsa karawang email:agung@unsika.ac.id nono heryana 2 study program system information faculty of computer science, universitas singaperbangsa karawang email: nono@unsika.ac.id ‹β› arip solehudin 3 study program technical information faculty of computer science, universitas singaperbangsa karawang email: arip.solehudin@staff.unsika.ac.id abstract—the progress of communication technology has had a positive impact on human life, including in the field of education. the education office is currently implementing computer-based exams starting from the state higher education entrance joint selection (sbmptn) exam to the school final examination. but with the implementation of computer-based exams this is of course the less secure the level, therefore the authors make this research with the aim of securing the exam data that will be tested with hill cipher cryptography and caesar cipher. cryptography is a technique of hiding data that is done to secure data, in this case cryptography aims to secure data on exam questions. keywords : kriptografi, hill cipher, caesar cipher abstrak — kemajuan teknologi komunikasi telah memberikan dampak positif pada kehidupan manusia, termasuk di bidang pendidikan. kantor pendidikan saat ini sedang melaksanakan ujian berbasis komputer mulai dari ujian seleksi bersama masuk perguruan tinggi negeri (sbmptn) hingga ujian akhir sekolah. tetapi dengan implementasi ujian berbasis komputer ini tentu saja semakin tidak aman levelnya, oleh karena itu penulis membuat penelitian ini dengan tujuan mengamankan data ujian yang akan diuji dengan kriptografi hill cipher dan cipher caesar. kriptografi adalah teknik menyembunyikan data yang dilakukan untuk mengamankan data, dalam hal ini kriptografi bertujuan untuk mengamankan data pada pertanyaan ujian. kata kunci: kriptografi, hill cipher, caesar cipher i. introduction currently technology has developed very rapidly, including in the field of education, an example of application technological development that is on the exam. the education office is currently implementing a computer-based exam starting from the joint higher education entrance examination (sbmptn) exams to the final school exams. but with the implementation of computer-based exams, of course the less the level of security, then of the authors make this research with the aim of securing the data on the exam questions to be tested with hill cipher and caesar cipher cryptography [1]. it can be interpreted that cryptography is hidden tulidan [2]. there are several algorithms or methods on cryptography includes hill cipher, caesar cipher, vernam cipher, advanced encryption standard (aes), and so forth. ii. method the method used in this study is the caesar cipher and hill cipher method for the process cryptography on data security exam questions, to combine the two methods there are several process that must be done. figure 2 shows the process carried out. fig 1 the cryptographic process of caesar cipher and hill cipher in figure 1 can be seen the cryptographic process of caesar cipher and hill cipher methods. the first process is the problem the test (plain text) is encrypted by the caesar cipher method and produces cipher text. then the cipher text converted to decimal. the result of decimal is re-encrypted using the hill cipher method generate cipher text [3].to process the description or return to the original message with the process carried out it is the opposite of the encryption process, if in the first encryption process with the caesar cipher method then on the first description process is the hill cipher method [4]. iii. results and discussion the cryptographic process using the hill cipher and caesar cipher method is done by entering an example exam 43 | vol.1 no.2, 10 july 2020 buana information tchnology and computer sciences (bit and cs) questions to be encrypted and make a key. first the message will be encrypted with using the caesar cipher method, then the results are converted to decimal and re-encrypted with the method hill cipher with a single process and a key [5]. examples of exam questions to be encrypted are english exam questions "i am so. i want to eat "with key (key) = 5. the first stage that will be done is encryption with the caesar cipher method. as for the process it is as follows: plaintext = i am so . i want to eat key = 5 then change the plaintext and key to binary data, can be seen in table 1. table i conversion from plaintext to binary plain text biner i 01001001 a 01100001 m 01101101 s 01110011 o 01101111 . 00101110 i 01001001 w 01110111 a 01100001 n 01101110 t 01110100 t 01110100 o 01101111 e 01100101 a 01100001 t 01110100 then do the encryption process by shifting the binary number 5 steps to the right, can be seen in table 2. table ii encryption process with caesar cipher biner cipher text 01001001 01001010 01100001 00001011 01101101 01101011 01110011 10011011 01101111 01111011 00101110 01110001 01001001 01001010 01110111 10111011 01100001 00001011 01101110 01110011 01110100 10100011 01110100 10100011 01101111 01111011 01100101 00101011 01100001 00001011 01110100 10100011 in table 2 the encryption results obtained by the caesar cipher method are still in the form of binary numbers. then do the conversion to decimal to facilitate the next encryption process. table iii binary to decimal conversion process cipher text decimal 01001010 74 00001011 11 01101011 107 10011011 155 01111011 123 01110001 113 01001010 74 10111011 187 00001011 11 01110011 115 10100011 163 10100011 163 01111011 123 00101011 43 00001011 11 10100011 163 in table 3 can be seen from the results of the conversion to decimal where the results will be directly encrypted with hill cipher method. in the hill cipher method, the key used is a matrix in which the matrix is used is 2x2 by using the same key in the encryption process with the method before that is 5 [6] [7]. so that the existing key can be used for the encryption process using the hill cipher method, the key will be formed 2x2 matrix by performing a simple calculation process [8]. key = 5 key k = [ 𝑘𝑒𝑦 𝑘𝑒𝑦 − 1 𝑘𝑒𝑦 + 1 𝑘𝑒𝑦 + 2 ] k = [ 5 5 − 1 5 + 1 5 + 2 ] so from the above calculation results obtained 2x2 matrix key with numbers k = [ 5 4 6 7 ] next divide the row of decimal numbers in ciphertext2 into a matrix with the number of key matrix columns (key = 2x2). [ 74 11 ] [ 107 155 ] [ 123 113 ] [ 74 187 ] [ 11 115 ] [ 163 163 ] [ 123 43 ] [ 11 163 ] then do the key matrix multiplication with the matrix that has been made. [5 4 6 7 ] [ 74 11 ] = [ 414 521 ] 𝑀𝑜𝑑 255 [159 11 ] [5 4 6 7 ] [ 107 155 ] = [1155 1727 ] 𝑀𝑜𝑑 255 [135 197 ] [5 4 6 7 ] [ 123 113 ] = [ 1067 1529 ] 𝑀𝑜𝑑 255 [ 47 254 ] [5 4 6 7 ] [ 74 187 ] = [ 1118 1753 ] 𝑀𝑜𝑑 255 [ 98 223 ] 44 | vol.1 no.2, 10 july 2020 buana information tchnology and computer sciences (bit and cs) [5 4 6 7 ] [ 11 115 ] = [515 871 ] 𝑀𝑜𝑑 255 [ 5 106 ] [5 4 6 7 ] [ 163 163 ] = [ 1467 2119 ] 𝑀𝑜𝑑 255 [ 192 79 ] [5 4 6 7 ] [ 123 43 ] = [ 787 1039 ] 𝑀𝑜𝑑 255 [ 22 19 ] [5 4 6 7 ] [ 11 163 ] = [ 707 1207 ] 𝑀𝑜𝑑 255 [ 197 187 ] in the last process of encryption change the results of this multiplication into the characters that can be seen in table 4. table iv the process of converting decimal to character decimal character 159 ÿ 11 vt 135 ‡ 197 å 47 / 254 þ 98 b 223 ß 5 enq 106 j 192 à 79 o 22 syn 19 dc3 197 å 187 » in table 4 we can see the final ciphertext results from the caesar cipher and hill cipher method, the ciphertext in the form of ascii numbers. so, the ciphertext from the example exam questions i am so. i want to eat is ÿ vt ‡ å / þ b ß enq j à o syn dc3 å ». furthermore, a description process is carried out to find out whether this method is successful for securing data on the sample exam questions. the first thing that will be done for the description process is the hill cipher method by multiplying the inverse key matrix with the ciphertext block matrix. k = [ 5 4 6 7 ] 𝑑𝑒𝑡k = (5 ∗ 7) − (4 ∗ 6) = 11 invers modulo: 11-1 mod 255 11x = 1 mod 255 11x = 1+255k x = (1+255k)/11 search for k = n with the result that x is an integer. k = 5; x = (1 + 255 * 5) / 11 = 116 (whole number) the inverse of 11 mod 255 is equivalent to 116 mod 255 which is 116. the determinant inverse modulo is used to find the martiks inverse. k = [ 5 4 6 7 ] then k-1 = determinan [ 𝑑 −𝑐 −𝑏 𝑎 ] so that k1 = 116 [ 7 −4 −6 5 ] = [ 812 −464 −696 580 ] 𝑚𝑜𝑑 255 = [ 47 46 69 70 ] continue multiplying the matrix with ciphertext. [ 47 46 69 70 ] [ 159 11 ] = [ 797 11741 ] 𝑀𝑜𝑑 255 [ 74 11 ] [ 47 46 69 70 ] [ 135 197 ] = [15407 23105 ] 𝑀𝑜𝑑 255 [ 107 197 ] [ 47 46 69 70 ] [ 47 254 ] = [ 13893 21023 ] 𝑀𝑜𝑑 255 [ 123 113 ] [ 47 46 69 70 ] [ 98 223 ] = [ 14864 22372 ] 𝑀𝑜𝑑 255 [ 74 187 ] [ 47 46 69 70 ] [ 5 106 ] = [5111 7765 ] 𝑀𝑜𝑑 255 [ 11 115 ] [ 47 46 69 70 ] [ 192 79 ] = [12658 18778 ] 𝑀𝑜𝑑 255 [ 163 163 ] [ 47 46 69 70 ] [ 22 19 ] = [ 1908 2848 ] 𝑀𝑜𝑑 255 [ 123 43 ] [ 47 46 69 70 ] [ 197 187 ] = [ 17861 26683 ] 𝑀𝑜𝑑 255 [ 11 163 ] after getting the results ,, then do the conversion to binary to proceed to the next process. tabel i proses konversi desimal ke biner decimal biner 74 01001010 11 00001011 107 01101011 155 10011011 123 01111011 113 01110001 74 01001010 187 10111011 11 00001011 115 01110011 163 10100011 163 10100011 123 01111011 43 00101011 11 00001011 163 10100011 these binary numbers are then re-encrypted for the last time using the caesar cipher method by shifting 5 times to the left, with the results that can be seen in table 6. table vi the process of converting decimal to binary 45 | vol.1 no.2, 10 july 2020 buana information tchnology and computer sciences (bit and cs) biner cipher text plain text 01001010 01001001 i 00001011 01100001 a 01101011 01101101 m 10011011 01110011 s 01111011 01101111 o 01110001 00101110 . 01001010 01001001 i 10111011 01110111 w 00001011 01100001 a 01110011 01101110 n 10100011 01110100 t 10100011 01110100 t 01111011 01101111 o 00101011 01100101 e 00001011 01100001 a 10100011 01110100 t in table 5 it can be seen that the results of the description that have been done produce a plaintext "iamso.iwanttoeat" in accordance with the example of the exam questions used for this study, with this the merging of the hill cipher and caesar cipher methods has been completed. iv. conclusion from the results of the research that has been carried out it can be concluded that the hill cipher and caesar cipher methods can be combined for the process of securing data with a good level of security, this method is also easily understood and for the encryption and description process using only one key so that it is easy to remember. references [1] r. kaur, “rectangular matrix with left inverse for variation in hill cipher: communication safe guard,” j. gujarat res. soc., vol. 21, no. 8, pp. 1234–1240, 2019. [2] c. h. bennett and g. brassard, “quantum cryptography: public key distribution and coin tossing,” arxiv prepr. arxiv2003.06557, 2020. [3] c. rajvir, s. satapathy, s. rajkumar, and l. ramanathan, “image encryption using modified elliptic curve cryptography and hill cipher,” in smart intelligent computing and applications, springer, 2020, pp. 675–683. [4] k. prasad and h. mahato, “cryptography using generalized fibonacci matrices with affine-hill cipher,” arxiv prepr. arxiv2003.11936, 2020. [5] i. irmayani, “application of matrix in hill cipher algorithm,” in international conference on natural and social sciences (iconss) proceeding series, 2019, pp. 141–147. [6] p. e. coggins iii and t. glatzer, “an algorithm for a matrix-based enigma encoder from a variation of the hill cipher as an application of 2× 2 matrices,” primus, vol. 30, no. 1, pp. 1–18, 2020. [7] d. kurniawan and b. priyatna, “pengamanan data berbasis mobile android dengan penggabungan linear feedback shift register (lfsr) dan modifikasi matriks kunci algoritma kriptografi playfair cipher,” j. telemat. mkom vol, vol. 10, no. 1, 2018. [8] a. behera, a. tripathy, a. r. tripathy, and s. rath, “random invertible key matrix decomposition for classical cryptography,” in advanced computing and intelligent engineering, springer, 2020, pp. 553–563. [9] s. kromodimoeljo, teori dan aplikasi kriptografi, jakarta: spk it consulting, 2010. [10] m. khoerudin, "algoritma hill cipher (sandi hill),"materi perkuliahan pada jurusan teknik informatika), 22 maret 2015. [11] sholeh, "algoritma subtitusi menggunakan chaesar cipher," caesar cipher dan cipher key, 3 oktober 2011. [12] k. w. m. r. puspita, "analisis kombinasi metode caesar cipher , vernam cipher , dan hill cipher dalam proses kriptografi," seminar nasional teknologi informasi dan multimedia 2015, pp. 43-48, 2015. [13] k. aulia, "soal ujian sekolah (us) bahasa inggris kelas 6 sd/mi tahun ajaran 2017/2018," juragan les, 2018. [14] irawan, a. s. y., el ramdhani, a. f., jordi, m., mahdi, r. s., & al mudzakir, t. (2020). implementasi algoritma advanced encryption standard (aes) untuk mengamankan file video. systematics, 2(1), 28-32. [15] hananto, a. l., solehudin, a., irawan, a. s. y., & priyatna, b. (2019). analyzing the kasiski method against vigenere cipher. arxiv preprint arxiv:1912.04519. p-issn : 2715-2448 | e-isssn : 2715-7199 vol.2 no.2 july 2021 buana information tchnology and computer sciences (bit and cs) issn : vol.xxx no.xxx 01 january 2020 44 | vol.2 no.2, july 2021 analysis of e-commerce adoption level on culinary micro, small and medium enterprises (umkm) in karawang regency using smart plus karya suhada 1 study program technical information stmik rosma, indonesia email: karya@rosma.ac.id lila setiyani2 study program information system stmik rosma, indonesia email: lila.setiyani@dosen.rosma.ac.id ‹β› damas setiadi sukardi 3 study program technical information stmik rosma, indonesia email: damas.sukardi@mhs.rosma.ac.id abstract— e-commerce is an electronic trading tool where trading transactions, both buying and selling, are carried out electronically on the internet network. the existence of the internet and various technologies in the telecommunication sector have changed many things, one of which is in umkm. the current owners of umkm are expected to be able to compete and maintain the continuity of their business by making changes and applications in the technical field. this study aims to analyze the level of e-commerce adoption in culinary umkm in karawang regency. the method used is a qualitative approach by measuring technology, environmental, organizational, and e-commerce adoption variables on the performance of umkm. the data collection technique used in this study was probability sampling, with random sampling types, with a total of 70 culinary umkm in karawang regency. the results of this study indicate that technology and environmental factors have a positive effect on the adoption of e-commerce so that they can improve the performance of the umkm in this study. keywords— e-commerce adoption, performance of umkm, umkm abstrak— e-commerce merupakan alat perdagangan elektronik dimana transaksi perdagangan, baik jual beli, dilakukan secara elektronik di jaringan internet. keberadaan internet dan berbagai teknologi di bidang telekomunikasi telah mengubah banyak hal, salah satunya di bidang umkm. para pemilik umkm saat ini diharapkan mampu bersaing dan menjaga kelangsungan usahanya dengan melakukan perubahan dan aplikasi di bidang teknis. penelitian ini bertujuan untuk menganalisis tingkat adopsi ecommerce pada umkm kuliner di kabupaten karawang. metode yang digunakan adalah pendekatan kualitatif dengan mengukur variabel adopsi teknologi, lingkungan, organisasi, dan e-commerce terhadap kinerja umkm. teknik pengumpulan data yang digunakan dalam penelitian ini adalah probability sampling, dengan jenis random sampling, dengan jumlah 70 umkm kuliner di kabupaten karawang. hasil penelitian ini menunjukkan bahwa faktor teknologi dan lingkungan berpengaruh positif terhadap adopsi e-commerce sehingga dapat meningkatkan kinerja umkm dalam penelitian ini. kata kunci— adopsi e-commerce, kinerja umkm, umkm i. introduction the development of technology at this time is very rapid, a business is required to utilize existing technology to run its business. one of the technological developments is in the business sector. currently, business owners are required to keep up with technological developments, because along with technological developments, traditional markets are being displaced by technology which in this case is ecommerce. according to kuswiratmo (2016:163) electronic commerce (e-commerce) or better known as online shopping is the implementation of commerce in the form of sales, purchases, orders, payments, and promotions of a product of goods and/or services carried out by utilizing computers and electronic communication facilities [1]. digital or data telecommunications. in addition, this form of commerce can also be carried out globally, namely by using the internet network [2]. it is undeniable that the development of information technology has had an impact on commercial activities. the existence of the internet and various technologies in the telecommunications sector have changed many things. umkm are expected to be able to compete and maintain business continuity by making changes and implementation in the technical field. according to the state ministry of cooperatives and small and medium enterprises (menegkop and ukm), what is meant by small business (uk), including micro enterprises (umi), is a business entity that has a net worth of at most rp. 200,000,000, excluding land and building a place of business, and having annual sales of a maximum of rp. 1,000,000,000. meanwhile, medium-sized enterprises (um) are business entities owned by indonesian citizens who have a net worth of more than rp. 200,000,000 to rp. 200,000,000. idr 10,000,000,000, excluding land and buildings [3]. karawang regency has a variety of umkm. in november 2020, there were 87,574 umkm actors registered in the karawang district [4]. by looking at the number of researchers interested in analyzing the level of adoption of ecommerce in kulimer umkm in karawang district to find out the benefits of e-commerce for umkm owners and to know the technological capabilities and levels of e-commerce mailto:lila.setiyani@dosen.rosma.ac.id mailto:damas.sukardi@mhs.rosma.ac.id 45 | vol.2 no.2, july 2021 adoption. the e-commerce referred to in this study is the use of go-food applications which are now increasingly being discussed as a means or activity of a business. the use of the go-food application is growing rapidly because in addition to making it easier for customers to find what food they want, it also makes it easier for business actors to make sales. the results of this study provide strategic knowledge and information regarding the adoption of e-commerce in culinary umkm in karawang district [5]. . ii. method a. types of research the research method used in this research is quantitative research. according to sugiyono (2011:8) [6], quantitative is a research method based on the philosophy of positivism, used to examine certain populations or samples, data collection using research instruments, data analysis is quantitative/statistical, with the aim of testing predetermined hypotheses. . meanwhile, according to borg and gall (1989) in quantitative research, causal relationships between variables are detected, then the observed information from the sample is obtained through statistical data collection. the data collected were analyzed in numerical form [7][8]. b research population according to suharsimi (1998:117) in [7] explained that the population is the whole object to be studied. meanwhile, according to sugiyono (2010) [9] the population is a generalization area consisting of objects or subjects that have certain characteristics and qualities that have been determined by a researcher to be studied and then conclusions will be drawn. the population in this study is umkm in the culinary field in karawang regency. iii. results and discussion the following is a recapitulation of the results of the questionnaire that researchers have distributed to culinary umkm in karawang regency, amounting to 75 respondents. table 1. recapitulation of questionnaire results variable code questionnaire answer sts ts n s ss technology (tn) tn1 0 1 10 21 43 tn2 0 1 12 25 37 tn3 0 3 17 23 32 tn4 2 1 17 22 33 organizational (or) or1 0 2 17 28 28 or2 1 3 20 27 24 or3 0 4 12 20 39 or4 0 1 15 17 42 environment (lk) lk1 1 2 10 27 35 lk2 2 4 17 26 26 lk3 0 0 15 25 35 lk4 2 4 18 26 25 e-commerce adoption (ac) ac1 0 0 5 23 47 ac2 0 1 8 30 36 ac3 1 0 9 23 42 ac4 1 1 14 35 24 ac5 1 0 13 33 28 ac6 0 1 6 23 45 ac7 2 5 23 23 22 ac8 0 4 17 27 27 performance umkm (kn) kn1 0 4 18 31 22 kn2 0 0 22 34 19 kn3 0 0 17 34 24 additional information : sts : strongly disagree ts : do not agree n : neutral s : agree ss : strongly agree data analysis results the data processing used in this research is using smartpls. in the results of this analysis, validity and reliability tests were carried out. the structural model in this study can be seen in the image below. fig.1 structural model a. validity test in the validity test, the indicator is considered valid if it has an outer loading value of the variable dimension having a loading value > 0.7 so it can be concluded that the measurement meets the criteria for convergent validity. the output generated by smartpls for outer loading is as follows [10]. table 2. outer loading results 46 | vol.2 no.2, july 2021 figure 2. outer loading value the results of the outer loading test in table 4.6 or figure 4.6 show that there are still invalid indicators marked in red. invalid indicators are ac4 with a value of 0.60, ac7 with a value of 0.654, ac8 with a value of 0.615 and tn4 with a value of 0.87. the condition for proceeding to the next stage is that the outer loading value must be valid, so that the outer loading test will be carried out again by removing / eliminating the previously invalid indicators. the results of the outer loading test can then be seen in the table below. b. reliability test the reliability test was carried out by looking at the composite reliability value. if the correlation value is more than 0.7 then it is said that the item provides a sufficient level of reliability, on the contrary if the correlation value is below 0.7 then the item is said to be less reliable. table 3. composite reliability value composite reliability e-commerce adoption 0.901 umkm performance 0.934 environment 0.894 organizational 0.788 technology 0.901 the table above shows that the correlation value of the composite reliability value for all constructs is above the value of 0.7. with the resulting value, all constructs have good reliability in accordance with the minimum value limit that has been required. to strengthen the results of the reliability test, the reliability test can be carried out with cronbach's alpha with the recommended value of 0.6. the table below shows that the cronbach's alpha value on all constructs has met the requirements, which are above 0.6. table 4. cronbach's alpha cronbach's alpha e-commerce adoption 0.862 umkm performance 0.912 environment 0.843 organizational 0.814 technology 0.833 c. evaluation of the structural model (inner model) structural model testing was conducted to see the relationship between the construct, the significance value and the r square of the research model. the value of r square is used to see the relationship between variables which is a goodnessfit model test. the value of r square can be seen. table 5. r square value r square r square ajusted e-commerce adoption 0.541 0.521 umkm performance 0.429 0.421 the table above shows that the e-commerce adoption variable has an r-square value of 0.541 which means that the technological, organizational and environmental variables affect the e-commerce adoption variable by 54.1% and the remaining 45.9% is influenced by other variables. while the umkm performance variable has an r-square value of 0.429, which means that the e-commerce adoption variable affects the umkm performance variable by 42.9% and the remaining 57.1% is influenced by other variables. d. t-statistic value the following is a diagram of the t-statistical values based on the output generated by smart pls. figure 3. output bootstraping adopsi ecommerce kinerja umkm lingkungan organisasional teknologi ac1 0.705 ac2 0.733 ac3 0.792 ac4 0.610 ac5 0.755 ac6 0.779 ac7 0.654 ac8 0.615 kn1 0.787 kn2 0.904 kn3 0.875 kn4 0.840 kn5 0.893 lk1 0.853 lk2 0.841 lk3 0.801 lk4 0.801 or1 0.840 or2 0.732 or3 0.881 or4 0.784 tn1 0.821 tn2 0.878 tn3 0.810 tn4 0.687 47 | vol.2 no.2, july 2021 iv. conclusion based on the results of the analysis that has been carried out, it can be concluded that the factors that influence the adoption of e-commerce umkm in the culinary field are technological, organizational and environmental factors. furthermore, to determine the relationship between variables, researchers tested the hypothesis. the results of this study indicate that : technological factors are proven to have a positive effect on e-commerce adoption. this shows that the availability of information technology tools, the availability of programs and support systems, the suitability of benefits and costs, as well as the capabilities and skills of human resources in using e-commerce can encourage umkm owners to use ecommerce in running their business. organizational factors are not proven to have a positive effect on e-commerce adoption. this shows that the availability of financial resources, the readiness of umkm owners to accept risks, leadership commitment and awareness of information technology developments do not encourage umkm owners to use e-commerce in running their business. environmental factors have been shown to have a positive effect on ecommerce adoption. this shows that the demands and encouragement from consumers, suppliers, the development of the business world, and competitive pressure are able to encourage umkm owners to use e-commerce in running their business. e-commerce adoption is proven to have a positive effect on umkm performance. this shows that ecommerce can facilitate access to information, improve business performance, improve quality and speed of service, improve cost efficiency, availability of facilities and infrastructure, keep up with technological developments, encouragement from external parties and support from all elements of e-commerce organizations can improve umkm performance. such as increased productivity, increased sales, increased profits/profits, increased product innovation and increased umkm innovation after using e-commerce. references [1] b. huda and b. priyatna, “penggunaan aplikasi content management system (cms) untuk pengembangan bisnis berbasis e-commerce,” systematics, vol. 1, no. 2, pp. 81–88, 2019. [2] helmalia, “pengaruh e-commerce terhadap peningkatan pendapatan usaha mikro kecil dan menengah di kota padang,” jebi (jurnal ekon. dan bisnis islam, vol. 3 no. 2, no. doi: 10.15548/jebi.v3i2.182, p. 237, 2018. [3] n. e. prastika and d. e. purnomo, “pengaruh sistem informasi akuntansi terhadap kinerja perusahaan pada usaha mikro kecil dan menengah (umkm) di kota pekalongan,” j. litbang, vol. 4, no. 3, 2019. [4] e. kurniawati, “87.574 umkm di karawang terdaftar menerima bantuanmodal,”tempo.co,p.available: https://metro.tempo.co/read/1404452/87-, 2020. [5] s. s. hilabi and b. priyatna, “pembangunan profil desa berkelanjutan sebagai wujud kuliah kerja nyata (kkn) berbasis online (studi kasus desa karawang kulon),” pros. konf. nas. penelit. dan pengabdi. univ. buana perjuangan karawang, vol. 1, no. 1, pp. 1732–1746, 2021. [6] sugiyono, metode penelitian kuantitatif, kualitatif dan r&d. bandung: alfabeta, 2011. [7] w. r. borg and m. d. gall, “educational research: an introduction, fifth edit,” new york and london: longman, 1989. [8] a. m. siregar, s. faisal, y. cahyana, and b. priyatna, “perbandingan algoritme klasifikasi untuk prediksi cuaca,” j. account. inf. syst., vol. 3, no. 1, pp. 15–24, 2020. [9] sugiyono, metode penelitian pendidikan pendekatan kuantitatif, kualitatif, dan r&d. bandung: alfabeta, 2010. [10] a. l. hananto and a. y. rahman, “user experience measurement on go-jek mobile app in malang city,” in 2018 third international conference on informatics and computing (icic), 2018, pp. 1–6. paper title (use style: paper title) issn : 2715-2448 | e-issn : 2715-7199 vol.2 no.1 january 2021 buana information technology and computer sciences (bit and cs) 11 | vol.2 no.1, january 2021 android based employee absence and leaving application information system baenil huda 1 information system, faculty of engineering and computer science universitas buana perjuangan karawang, indonesia baenil88@ubpkarawang.ac.id shofa shofiah hilabi 2 information system, faculty of engineering and computer science universitas buana perjuangan karawang, indonesia shofa.hilabi@ubpkarawang.ac.id ‹β› maya rahayuningsih 3 information system, faculty of engineering and computer science universitas buana perjuangan karawang, indonesia si16.mayarahayuningsih@mhs.ubpkarawang.ac.id abstract —attendance in general is the recording of employee attendance and is one aspect of assessment in a company. the purposes of this study are finding out how employees can apply for leave and how to design and build an android-based attendance system application. data collection methods that used in this research include observation, interviews and literature study. the system development method that will be used is the system development life cycle (sdlc) model of the waterfall. with this system, employees and companies will be helped in the problem of absence and leave. the company will get accurate, fast and precise data in decision making. keywords: attendance, leave, android, uml, waterfall abstrak — absensi secara umum merupakan pencatatan kehadiran karyawan dan termasuk salah satu aspek penilaian dalam suatu perusahaan. tujuan dari penelitian ini yaitu pertama untuk mengetahui bagaimana karyawan dapat melakukan pengajuan cuti dan yang kedua yaitu untuk mengetahui bagaimana merancang dan membangun aplikasi sistem absensi berbasis android. metode pengumpulan data yang akan digunakan pada penelitian ini yaitu diantaranya observasi, wawancara dan studi pustaka. metode pengembangan sistem yang akan digunakan yaitu system development life cycle (sdlc) model air terjun (waterfall). dengan adanya sistem ini, karyawan dan perusahaan akan terbantu dalam masalah absen dan cuti serta perusahaan akan mendapatkan data yang akurat, cepat dan tepat dalam pngambilan keputusan. kata kunci: absensi, cuti, android, uml, waterfall i. introduction the development of technology as it increases from year to year is now very fast, we can get various technological advances easily, especially in information technology. utilizing information technology is a must for companies so that they are not left behind by the times, so, naturally, today many companies are competing to update new systems and technologies because to improve their business, especially in fields that are closely related to information technology. one of them is the employee attendance system and leave requests, the company must carry out this process properly, to make it easier to make decisions, the company needs accurate information without having to go through the manual recording which must be done repeatedly because the manual process takes a long time so it is not effective. . previously, the company only had physical documents but over the years the company-owned data also increased. this is one of the reasons companies use information technology in managing data such as attendance and leave data. attendance systems at companies start from using paper, fingerprints, magnetic cards to only using smartphones. the number of smartphone users currently allows some companies to update their android-based attendance and leave application systems [1] because it is more effective and efficient without having to queue to do absences and to reduce the buildup of leave application files in general. online attendance is recording attendance with a system that is connected to a real-time database [2]. employees can take attendance via smartphone anywhere according to their entry and return hours as long as the employee is still in the company environment. this requires a local area network (lan) which is only within the company environment so that employees cannot do attendance outside the company environment because the network coverage that is set up only covers company areas. besides being able to do attendance, this system can also apply for leave. the concept of an attendance system application that will be made is to make it easier for employees and companies in making attendance and filing leave and in making reports related to company data [3]. based on the above background, the writer tries to analyze and study the attendance system and leave submission which i will write in a final report entitled "application of employee attendance information system and androidbased leave submission". ii. methods mailto:baenil88@ubpkarawang.ac 12 | vol.2 no.1, january 2021 in this study, researchers researched several stages, namely the first to collect data, design systems, develop systems, test, and write a final report. the stages are as follows [4]: teknik pengumpulan data pengembangan sistem hasil penelitian waterfall metode penelitian observasi studi pustaka wawancara analysis design pengodean testing maintenance figure 1. research flowchart based on the picture above, the research stages are carried out starting from collecting data about the object to be studied, then continuing with designing the system to be made using the unified modeling language (uml) after that developing the system using the waterfall model [5], then proceeding to test a system to check functionality, if all stages have been completed then the process of writing a report from the research results. systems development method the system development method used in this study is the waterfall methodology [6]. figure 2.stages of the waterfall methodology a. analysis the process of collecting requirements is carried out intensively to specify software requirements so that it is easy to understand what kind of software is needed by the user [7]. in this study, the analysis was carried out through the method of observation at pt. xyz regarding attendance and leave as well as conducting direct interviews with several employees of pt. xyz to determine the shortcomings of the ongoing system and researchers dig up data from several journals, theories, and electronic documents that can support the research process [8]. b. design a multi-step process that focuses on the design of a software program including data structures, interface representation software architectures, and coding procedures. in this study, researchers designed a system using uml (unified modeling language) and for user interface design using the pencil application [9]. making program code the design must be translated into a software program. the result of this stage is a computer program by the designs that have been made at the design stage. in making program code (coding) in this study using android studio tools, with the flutter framework and the dart programming language and visual studio code as a programming language code editor. c. making program code the design must be translated into a software program. the result of this stage is a computer program by the designs that have been made at the design stage. in making program code (coding) in this study using android studio tools, with the flutter framework and the dart programming language and visual studio code as a programming language code editor. d. maintenance (maintenance), the maintenance stage can repeat the development process starting from specification analysis for changes to existing software, but not for creating new software. maintenance includes correction of unknown errors in the previous process, improvement of implementation, development of the system unit, and program maintenance. iii. results and discussion in making program code (coding) in this study using android studio tools, with the flutter framework and the dart programming language and visual studio code as a programming language code editor. before entering the coding stage, make a design and workflow of the proposed system first. the following are some of the stages in system and software design [10]: 1. use case diagram. use case diagrams are pictures of some or all actors and use cases to recognize their interactions in a system a. use case diagram of employee attendance system 13 | vol.2 no.1, january 2021 figure 1. usecase diagram of employee attendance system 2. activity diagram activity diagrams describe a series of flow from activities, used for other activities such as use cases or interactions. a. employee attendance activity diagram gambar 1. activity diagram absensi karyawan to carry out the employee attendance process, you must first log in using the username and password that has been set by the admin for each employee, if the verification is successful it will go to the main page and if it fails, it will return to the login page. when you have successfully logged in, select the attendance menu then select checkin when entering or select check out when leaving. b. activity diagram for requesting leave figure 3. activity diagram for requesting leave that is, to apply for leave such as the activity above, the employee must first log in first, then if the login is successful, it will enter the main page, if it fails, the login will return to the login menu when you have successfully logged in, select the left menu then see the remaining leave then fill in the leave submission form, after that the leader approves the leave submission, the admin confirms and the employee waits for confirmation from the admin whether the leave application is accepted or rejected by looking at the leave status in the leave history menu then print leave form. 3. sequence diagram sequence diagrams describe the dynamic collaboration between some objects and to show a series of messages sent between objects as well as interactions between objects, something that occurs at a certain point in the system's execution. a. employee attendance sequence figure 4. sequence diagram of employee attendance employe attendance system, namely to enter the application, the employee must first enter a karyawan aplikasi leader admin login verifikasi halaman utamapilih menu cuti tampil sisa cuti isi form cuti menunggu konfirmasi ya tidak cetak approval konfirmasi 14 | vol.2 no.1, january 2021 username and password. if the login is successful, it will enter the main page, if it fails, it will return to the login page. when you have successfully logged in, select the attendance menu, then the system will display an attendance page after that select check-in when entering or check out when exiting and select log out to exit the application. b. sequence diagram for filing leave figure 5. sequence diagram for filing leave to apply for leave, employees must first log in by entering the username and password that has been set by the admin for each employee, if the employee successfully logs in, he will enter the main page, and if it fails, it will return to the login page. when login is successful, select the left menu then the system will display the left page then see the remaining leave then fill in the leave application form and it is approved by the leader after that wait for confirmation from the admin whether the application is accepted or rejected via the application then print the leave form that has been received. select logout to exit the application. 4. class diagram class diagrams are used to explain the system structure in terms of defining the classes that will be made to build a system. a. class diagram of employee attendance system and leave application figure 6. class diagram of the employee attendance system and leave requests designing application and web pages for employee attendance and leave submission page design interface design is an interface design that describes the display plan of the employee attendance system application to be built. below is a view of several interface designs for mobile and web admin applications made using the pencil application [11]. figure 9. design of employee attendance page on this attendance design page, employees perform attendance by choosing check-in for entry and check-out on exit figure 10. design of leave application page on this leave design page, employees apply for leave only by filling in the leave form and then the leader will approve it 15 | vol.2 no.1, january 2021 figure 11. the design of the admin attendance page on the attendance data admin web design page, the admin can view the attendance data of all employees and download it. figure 12. the design of the leave leader's approval page on the leader's leave approval design page, the leader can approve the employee leave application, which will then be confirmed by the admin. figure 13. the design of the admin leave confirmation page on this web admin design page, the admin can confirm submission of employee leave after leader approval. implementation of employee attendance and leave application web pages the results of the interface design have been built in the first stage, then the next stage is implementing or realizing the system design into the actual application. below are some of the display implementations of the mobile application and the web admin for employee attendance systems. figure 14. implementation of employee attendance pages employee attendance page where employees only click checkin when entering or check out when leaving then the data will be stored in the database, and employees can view attendance history or can print it if needed [12]. figure 15. implementation of leave application page the leave application page is for employees to apply for leave by simply filling in the form and entering the date from and to when the employee will apply and employees must also fill in family contacts who can be contacted during leave [13]. figure 16. implementation of the main admin page 16 | vol.2 no.1, january 2021 attendance system dashboard page which is the main web page and can only be accessed by admin [14], on that page can be seen the number of employees, attendance, and leave requests. figure 17. implementation of admin attendance page attendance page where admin can see the attendance data of all employees and print it to make a report that will be submitted to superiors. figure 18. implementation of admin leave confirmation page leave confirmation page on the admin web where admin confirms employee to leave after the leader approves and the application will receive a notification if the leave request has been validated by the admin [15]. iv. conclusion based on the results of the above discussion and observations, several conclusions can be drawn, namely: 1. employees can do attendance without queuing using the online attendance application with the username and password that the admin has set for each employee. 2. employees can apply for leave online by simply filling in the leave form through the application. references [1] p. a. sunarya, e. febriyanto, and j. januarini, “aplikasi mobile absensi karyawan dan pengajuan cuti berbasis gps,” ccit j., vol. 12, no. 2, pp. 241– 247, 2019. [2] r. a. makhfuddin and n. prabowo, “aplikasi absensi menggunakan metode lock gps dengan android di pln app malang basecamp mojokerto,” issn, vol. 5, no. 2, pp. 55–63, 2015. [3] dhanta dikutip dari sanjaya, “aplikasi berbasis web,” aplikasi berbasis web. 2015. [4] s. mulyani, metode analisis dan perancangan sistem. 2017. [5] d. m.shalahuddin, “model pengembangan perangkat lunak menurut rosa a . s . dan m . shalahuddin,” model pengemb. perangkat lunak, 2014. [6] s. s. hilabi and . p., “analisis kepuasan pengguna terhadap layanan aplikasi media sosial whatsapp mobile online,” buana ilmu, vol. 3, no. 1, pp. 119– 136, 2018. [7] m. s. rosa a.s, rosa a.s, m. s. (2016). model waterfall. 2016. bandung: informatika.model waterfall. 2016. [8] m. y. simargolang and w. a. warsito, “analisis sistem pengolahan absensi karyawan pada pt. bakrie sumatera plantations tbk bunut,” j. teknol. inf., vol. 1, no. 2, p. 114, 2018. [9] rosa a.s dan m. shalahudin, “rekayasa perangkat lunak (terstruktur & berorientasi objek),” politek. negri sriwij., 2011. [10] verdi yasin, rekayasa perangkat lunak berorientasi objek. 2011. [11] baenil huda and saepul apriyanto, “aplikasi sistem informasi lowongan pekerjaan berbasis android dan web monitoring (penelitian dilakukan di kab. karawang),” buana ilmu, 2019. [12] s. n. aisyah and k. a. hafizd, “aplikasi absensi karyawan pt. angkasa pura i (persero) banjarmasin,” j. sains dan inform., vol. 3, no. 1, p. 7, 2017. [13] a. setiyanto, f. samopa, and alwi, “pembuatan sistem informasi cuti pada kantor pelayanan perbendaharaan negara dengan menggunakan php dan mysql,” tek. pomits, vol. 2, no. 2, pp. 381–384, 2013. [14] b. huda, “sistem informasi data penduduk berbasis android dan web monitoring studi kasus pemerintah kota karawang (penelitian dilakukan di kab. karawang),” buana ilmu, 2018. [15] b. priyatna, s.s hilabi, n. heryana, “aplikasi pengenalan tarian dan lagu tradisional indonesia berbasis multimedia,” systematic, vol. 1, no. 2, pp. 89–98, 2019. paper title (use style: paper title) p-issn : 2715-2448 | e-issn : 2715-7199 vol.1 no.2, 10 july 2020 buana information tchnology and computer sciences (bit and cs) 37 | vol.1 no.2, 10 july 2020 buana information tchnology and computer sciences (bit and cs) design of reminder information last contract (rilc) applications using web-based sms gateway arip solehudin 1 program study teknik informatika fakultas ilmu komputer, universitas singaperbangsa karawang arip.solehudin@staff.unsika.ac.id aldi angga putra 2 program study teknik informatika fakultas ilmu komputer, universitas singaperbangsa karawang aldi.angga@student.unsika.ac.id kamal prihandani 3 program study teknik informatika fakultas ilmu komputer, universitas singaperbangsa karawang kamal.prihandani@unsika.ac.id nono heryana 4 program study sistem informasi fakultas ilmu komputer, universitas singaperbangsa karawang nono@unsika.ac.id abstract— currently informarsi system is needed by several large companies. to support some data processing, one of them is employee data processing and employee contract data. given the existing problems in contract data pengelohan often occurs that the check is not maximal so some data left behind and even neglected impact on the increase in salaries of employees who are late due to contracts that are not handled according to schedule. this study aims to create a web-based ril application using sms gateway for delivery of contract data reminder information. this research uses sdlc method with extreme programming model with stages of planning, designing, coding, testing, implementation and evaluation. sms gateway is a gateway for information dissemination using sms. with sms gateway can spread the message of the serial number automatically and quickly which is directly connected with the database of mobile numbers that have been stored. this can be done by ril applications in sending reminder information messages to the hrd division as well as employees receiving their contract expiration date information. the result of the research is the user can maximize the checking of contract data so that nothing is left behind and give the report that arranged well and clear. keywords— contract data management, sms gateway, web abstrak— saat ini sistem informarsi dibutuhkan oleh beberapa perusahaan besar. untuk mendukung beberapa pemrosesan data, salah satunya adalah pengolahan data karyawan dan data kontrak karyawan. mengingat masalah yang ada dalam pengelohan data kontrak sering terjadi cek yang tidak maksimal sehingga beberapa data tertinggal dan bahkan terabaikan berdampak pada kenaikan gaji karyawan yang terlambat karena kontrak yang tidak ditangani sesuai jadwal. penelitian ini bertujuan untuk membuat aplikasi ril berbasis web menggunakan sms gateway untuk pengiriman informasi pengingat data kontrak. penelitian ini menggunakan metode sdlc dengan model extreme programming dengan tahapan perencanaan, perancangan, pengkodean, pengujian, implementasi dan evaluasi. sms gateway adalah gateway untuk penyebaran informasi menggunakan sms. dengan sms gateway dapat menyebarkan pesan nomor seri secara otomatis dan cepat yang terhubung langsung dengan database nomor ponsel yang telah disimpan. ini dapat dilakukan oleh aplikasi ril dalam mengirimkan pesan informasi pengingat ke divisi hrd serta karyawan yang menerima informasi tanggal kedaluwarsa kontrak mereka. hasil dari penelitian ini adalah pengguna dapat memaksimalkan pemeriksaan data kontrak sehingga tidak ada yang tertinggal dan memberikan laporan yang tertata dengan baik dan jelas. kata kunci— manajemen data kontrak, sms gateway, web i. introduction the company has important assets, one of the most important assets of the company in addition to data, namely the managing human resources is the hrd (human resources development) division. one of his duties is to deal with employee contract issues. for that reason, in providing information about employee contracts that will be implemented usually use the short message service (sms). sms is one of the most popular cellular services at the moment. for companies and agencies, sms gateway is needed because sms gateway can provide information facilities related to the activities of the company or agency. in practice, the processing of employee contract data still uses microsoft excel as the main archive material which still experiences several obstacles namely difficulties in checking employee contracts, often even one of the employees who should have finished the contract or the contract will be ignored. with the large number of employee contract data that must be processed and increasingly complex, problems must be addressed and the need for information precisely and quickly. even the processing of employee contract data does not have an accurate report. then every time a contract is carried out suddenly without being prepared in advance because it has not been maximized in monitoring employee contract data, which results in fatal salary for employees who should rise when a new contract, but instead becomes left behind even neglected. therefore, it is necessary to design a web-based system and contract reminders using an sms gateway so that it can provide information more quickly and accurately. in addition, monitoring of information, especially employee contract data, is no longer overlooked or left behind in the contract period. this system also provides a reminder message facility via sms when employee contract data will be exhausted or not yet contracted and will soon be contracted. mailto:arip.solehudin@staff.unsika.ac.id mailto:arip.solehudin@staff.unsika.ac.id mailto:arip.solehudin@staff.unsika.ac.id mailto:aldi.angga@student.unsika.ac.id mailto:kamal.prihandani@unsika.ac.id mailto:nono@unsika.ac.id 38 | vol.1 no.2, 10 july 2020 buana information tchnology and computer sciences (bit and cs) based on the description above, for this reason, this research builds an application "design of reminder information last contract (ril) application using webbased sms gateway" by using the extreme programming (xp) model sdlc methodology including the stages carried out including: planning, designing, coding and expected testing can provide convenience in the delivery of information needed both management or employees themselves so that the employee contract process according to the contract expiration date. ii. method the methodology used is the sdlc method with the xp model. which can reduce development costs (implementation phase), a semi-formal methodology. planning developers must always be ready for changes because changes will always be accepted, or in other words flexible (maintenance phase), there are four stages of system development in this xp model, namely: planning, design, coding, and testing. figure 1 research flow iii. results and discussion the results of the research discussed in the previous chapter, will be discussed in this chapter. in this study the design of a reminder information last contract (ril) application using a web-based sms gateway with the xp model consists of several stages including planning, user story, design, coding, and testing. 3.1 planning planning is done to find out the current problems and to find out the needs needed for the application that will be made in this study. 3.1.1 user stories based on data collection by means of interviews obtained problems experienced by potential users as follows, 1. managing irregular employee contract data. 2. employee contract data checking is only done by looking at employee contract data in microsoft excel to see the expiration date of the employee contract so that sometimes the data is often missed plus there is no reminder information in checking the employee contract data. 3. inaccurate employee contract data reports. there are always differences in employee contracts, one of which is salary and the expiration date of the employee contract. 4. don't have an employee contract data processing application as well as an information reminder about employee contract data. from the problems mentioned above there are several application needs to know clearly and precisely the current problems faced by potential users of this application. 1. functional needs functional requirements analysis is carried out to find out what needs are needed by users of the reminder information last contract (ril) application using a web-based sms gateway. from the results of the interview in accordance with the user stories that have been made, then obtained some of the needs of users of the reminder information last contract (ril) application using web-based sms gateway. a. assist in declining employee contract data. b. helps check and provide employee contract data reminder information. c. provide clear and accurate reports. 2. system user analysis analysis of system users is done to find the right user for the application to be made. the analysis results obtained from the classification of system users, this application can only be used by the hrd division who knows all employee data, one of which is employee contract data. whereas the employees themselves only receive a piece of information about the continuation of their contracts, and the director himself only receives a contract report from the system user, hrd. 3.2 design 3.2.1 design modeling the design of this application will be built using uml modeling. where the application architecture design is created using class responsibility collaborator (crc), use case diagrams, activity diagrams, and sequence diagrams. 1. class responsibility collaborator (crc) crc is a collection of standard index cards that have been divided into three parts (classes, responsibilities, collaborators) 39 | vol.1 no.2, 10 july 2020 buana information tchnology and computer sciences (bit and cs) figure 2 crc employee contract data 2. use case diagrams the use case diagram describes an interaction between theactor and the application that will be created as shown below: figure 3. use case diagram of ril applications 3. activity diagram activity diagram explains the activity flow of the system process, with the activity diagram showing more detailed flow of the system running sequentially. figure 4 contract data check activity diagram 4. sequence diagram sequence diagram illustrates the interaction between objects around the system. figure 5. sequence diagram check contract data 5. class diagram class diagrams are used to describe the types of objects in a system, class diagrams also show the properties and operations of a class and the constraints that exist in relation to an object. the following figure 6 is a class diagram presentation from the rimender information last contract (ril) application. figure 6. class diagram of ril applications 3.2.2 designing design 1. interface design in the interface design a menu structure will be made which is shown below: a. employee contract data layout next figure 7 is a display of employee contract data from the ril application. system data basedivisi hrd login berhasil halaman utama mengecek data kontrak data kontrak karyawan habis kontrak <= 1 minggutidak menemukan data mengirimkan pesan ke pihak hrd menemukan data dapat verifikasi mengirimkan pesan ke karyawan untuk tanda tangan kontrak diperpanjang kontrak kirim pesan ke karyawan informasi tidak diperpanjang kontraknya tidak diperpanjang login +username +password +login() +cancel() +lupa password() data user +username +password +no.telp +tambah() +edit() +hapus() +cek data() data karyawan +nik +nama lengkap +jabatan +tempat lahir +tanggal lahir +alamat +no.kk +no.ktp +no.telp +tambah() +edit() +hapus() data kontrak +nik +tanggal habis kontrak +tanggal tanda tangan +gaji yang di terima +kontrak ke +tambah() +edit() +hapus() laporan +nik +nama +jabatan +tanggal habis kontrak +tanggal tanda tangan +gaji yang diterima +kontrak ke +cetak() +unduh() 40 | vol.1 no.2, 10 july 2020 buana information tchnology and computer sciences (bit and cs) figure 7. employee contract data b. report layout next figure 8 is a report display from the ril application. figure 8. report 3.3 coding from the design that has been made then will be translated into a software program (software) whose result is a web-based information last contract (ril) reminder application created using html, css, java script, php and phpmyadmin program codes as its database. the following application views for reminder information last contract (ril): a. display contract data menu this display will appear if the user selects the contract data menu in the main menu and in this menu there is an application contract data table from real that we can add, edit or delete and check the contract data that is exhausted. the following picture 9 display contract data. figure 9. contract data display b. report menu display this display will appear if the user selects the report menu in the main menu and in this menu there is a contract data report table that we can save or print. the following figure 10 display contract reports. figure 10. report views 3.4 testing at this stage the user (user) tries an application that has been built according to the user's request. the aim of the user is to evaluate this system in order to see what the user wants as in the user story design stage and identify the problems that occur in the application even though it has been tested before. evaluation is done by conducting interviews with the hrd manager of pt. fajar putra nusantara karawang: 1. can the ril application manage contract data easily? answer: yes, employee data management can be accessed and managed anywhere ... 2. can the ril application help with checking contract data? answer: yes, it is very helpful so as to minimize contract data that is left behind. 3. can the ril application help in remembering contract data? answer: yes, it is very helpful in remembering good contract data for our hrd division and even the employees concerned. 4. can the ril application provide clear and accurate reports? answer: yes, the report is quite clear and accurate. 5. does the function of the ril application fit the needs? answer: yes, the functional needs of the application are sufficient and are as needed. 6. is the ril application easy to understand? answer: yes, this application is easy to understand. iv. conclusion based on the stages of the research that has been done in making this reminder information last contract (ril) application, the following conclusions can be drawn: 1. ril application is applied at pt. fajar putra nusantara to maximize the performance of the hrd division in contract data management. 2. the ril application utilizes sms gateway technology in providing in remittance information for contract data that will run out.. references [1] afrina, m., & ibrahim, a. (2015). pengembangan sistem informasi sms gateway dalam meningkatkan layanan komunikasi sekitar akademika fakultas ilmu komputer unsri. jurnal, 855-856. [2] edwin b, j., & widayati, s. (2013). jurnal. aplikasi pencarian informasi surat tanda nomor kendaraan (stnk) berbasis sms gateway, 115. 41 | vol.1 no.2, 10 july 2020 buana information tchnology and computer sciences (bit and cs) [3] ahmad, s. b., & amir, r. (2020). sistem kontrak kerja antara karyawan dan perusahaan perspektif undang-undang ketenagakerjaan dan hukum islam (studi kasus di pt citra van titipan kilat). shautuna: jurnal ilmiyah mahasiswa perbandingan mazhab dan hukum, 1(2). [4] handayani, i., aini, q., & oktavyanti, y. (2015). penggunaan rinfocal sebagai aplikasi pengingat. jurnal, 3. [5] d. juardi, “presensi dan reminder menggunakan qr code (studi kasus : sma xxx),” systematics, vol. 1. pp. 33–43, 01-aug2019. [6] nurhayati. (2015). aplikasi sms reminder pada perpustakaan apikesakbid citra medika surakarta. jurnal, 40. [7] adam, a., & amri, h. (2019). prototype monitoring arus dan tegangan menggunakan sms gateway. multitek indonesia, 13(1), 1623. [8] nurlaela, f. (2013). aplikasi sms gateway sebagai sarana penunjang informasi perpustakaan pada sekolah menengah pertama negeri 1 arjosari. indonesian journal on networking and security, 20-25. [9] solehudin, a., & garno, g. (2017). prototype api pada aplikasi pembatasan akses internet dengan pemanfaatan hak akses user profile hotspot. jurnal rekayasa informasi, 6(2). [10] arifin, n. (2019). analisa risiko kontrak kerja lumpsum pada proyek gedung k3 surabaya. axial: jurnal rekayasa dan manajemen konstruksi, 7(1), 25-32. [11] fathoni, k., fariza, a., & firmansyah, y. e. (2019). rancang bangun sistem informasi manajemen kegiatan penelitian di politeknik elektronika negeri surabaya. jurnal ilmiah teknologi informasi asia, 14(1), 7-18. [12] solehudin, arip, nono heryana, and yana cahyana. "designing and building client-server based student admission applications." buana information technology and computer sciences (bit and cs) 1.1 (2020): 16-18. [13] hananto, a. l., solehudin, a., irawan, a. s. y., & priyatna, b. analyzing the kasiki method against vigenere cipher. [14] nuraeni, a., & hardianti, s. s. (2019). aplikasi penerimaan karyawan online dengan fitur informasi jadwal tes dan hasil kelulusan. internal (information system journal), 2(1), 1-22. [15] r. mayasari, “sistem informasi nilai mahasiswa berbasis sms gateway menggunakan trigger pada database,” systematics, vol. 1. pp. 44–57, 01-aug-2019. paper title (use style: paper title) issn : 2715-2448 | e-issn : 2715-7199 vol.2 no.1 january 2021 buana information technology and computer sciences (bit and cs) 1 | vol.2 no.1, january 2021 buana information technology and computer sciences (bit and cs) implementation of simple additive weighting (saw) method to determine exemplary pkh social worker (case study: ppkh garut regency) topan setiawan information systems study program faculty of computer science universitas ma’soem bandung, indonesia topansetiawan@masoemuniversity.ac.id bayu priyatna information system, faculty of engineering and computer science universitas buana perjuangan karawang, indonesia bayu.priyatna@ubpkarawang.ac.id ‹β› abstract—one measure of the success of an employee in carrying out his job is his election as an exemplary employee in the agency where he works. the selection must of course be based on standard measurement and objective assessment, the goal is that the predetermined results can be justified. a decision support system using the simple additive weighting (saw) method can help ppkh garut regency in determining exemplary pkh social worker in each sub-district ppkh unit, this method will look for weight values for each predetermined criterion consisting of quantity, integrity, dedication, reliability, initiative, diligence, attitude, motivation and presence. the data sample taken in this research was ppkh x sub-distric, where the results of the research are in the form of a ranking that can support the decision to choose an exemplary pkh social worker in ppkh garut regency. keywords: decision support system, pkh social worker, ppkh garut regency. abstrak —salah satu ukuran keberhasilan seorang pegawai dalam menjalankan pekerjaannya adalah terpilihnya sebagai pegawai teladan di instansi tempatnya bekerja. pemilihan tersebut tentunya harus didasarkan pada standar ukuran dan penilaian yang objektif, tujuannya agar hasil yang telah ditentukan dapat dipertanggung jawabkan. sistem pendukung keputusan dengan menggunakan metode simple additive weighting (saw) dapat membantu ppkh kab. garut dalam menentukan pendamping sosial pkh teladan di setiap unit ppkh kecamatan, metode ini akan mencari nilai bobot untuk setiap kriteria yang telah ditentukan yang terdiri dari kuantitas, integritas, dedikasi, kehandalan, inisiatif, kerajinan, sikap, motivasi dan kehadiran. sampel data yang diambil pada penelitian ini adalah ppkh kec. x, dimana hasil penelitian berupa pemeringkatan yang dapat mendukung keputusan pemilihan pendamping sosial pkh teladan di ppkh kab. garut. kata kunci: sistem pendukung keputusan, pendamping sosial pkh, ppkh kab. i. introduction program keluarga harapan (pkh) is a program that provides conditional social assistance to keluarga penerima manfaat (kpm) that has been established by the ministry of social affairs or kementerian sosial (kemensos), as an effort to accelerate poverty reduction and has been launched by the government of indonesia since 2007 [1]. pkh implementers or pelaksana pkh (ppkh) are agencies scattered in every region starting from the central level (ppkh pusat), provincial level (ppkh provinsi), regency or city level (ppkh kabupaten / kota) and subdistrict level (ppkh kecamatan). pkh social worker at the sub-district level are employees or people who have direct contact with kpms to ensure that kpms get their rights and carry out their obligations in accordance with the terms and conditions. in an uncertain period of time, ppkh garut regency in collaboration with the garut regency social service has twice selected exemplary pkh social workers to give appreciation to workers who have shown achievements in carrying out their work. the selection was carried out on a bottom-up basis starting at the sub-district level where workers selected from the sub-district level were submitted to the regency level to take further tests. in its implementation the standards for measuring and assessing work performance are less clear and not transparent, so that dissatisfaction for those who feel that their performance has been maximized but are not selected. this in turn creates a negative stigma that the selection is subjective and the reward received is not based on work performance, but on other factors outside of work assignments. to overcome this problem, it is necessary to implement a decision support system with standardized measures so that the results can be accounted for [1]. decision support system (dss) is a system that is used to assist and determine decisions to information users to be more precise in solving problems that exist within a company, agency, or organization by data and certain methods [2]. one method that can be applied is the simple additive weighting (saw) method where the result is a ranking of exemplary employees [3]. by using the dss in decision making, the standard for measuring and evaluating the performance of each worker will be determined properly, so that all parties involved can accept the decision [2]. mailto:topansetiawan@masoemuniversity.ac.id mailto:bayu.priyatna@ubpkarawang.ac.id 2 | vol.2 no.1, january 2021 ii. method a. research framework input output study of literature data collection data processing reporting understanding theories & concepts data and information required list of worker ratings research report fig. 1. research framework b. saw method saw is a simple multi-criteria decision-making method [4]. the steps in this method include: 1. assessment criteria (cj, j = 1, 2, 3,…, m), which are used as a reference in making this decision are shown in table i. table i. assessment criteria kode criteria (cj) c1 quantity c2 integrity c3 dedication c4 realibility c5 initiative c6 diligence c7 attitude c8 motivation c9 presence information: quantity : how quickly the worker gets the job done integrity : how committed the worker is to the job dedication : how much dedication is devoted by worker to realize the ideals and success of the pkh program realibility : relates to whether or not worker can be relied on on certain issues initiative : how often worker take corrective action, make suggestions for job improvement and accept responsibility for completing work diligence : willingness to carry out tasks without coercion and also of a routine nature attitude : worker's behavior towards the organization or boss or coworkers motivation : how successful are worker in motivating kpm to leave pkh program participation. presence : how often are worker present at the workplace to work or attend internal organization meetings 2. determine the weight of each criterion with (wj, j = 1, 2, 3, …, m) where ∑wj = 1. 3. determine a decision matrix using equation (1). if j is benefit criteria (1) if j is cost criteria information: rij : normalized performance rating value xij : the attribute value that each criterion has max xij : the largest value of each criterion min xij : the smallest value of each criterion benefit : if the largest value is the best cost : if the smallest value is the best 4. calculating the preference value for each alternative using equation (2). (2) information: vi : ranking for each alternative wj : the weighted value of each criterion rij : normalized performance rating value the biggest vi value indicates that the alternative ai is an alternative choice. iii. results and discussion a. determination of the scale and weight of the criteria determination of alternative values for each criterion using a likert scale of 9-1 [5], with the following formulations: table ii. criteria likert scale value quality 9 a 8 a 7 b+ 6 b 5 b 4 c+ 3 c 2 c 1 d the weight for each criterion based on the value of importance is addressed as in table iii. tabel iii. criteria weight 3 | vol.2 no.1, january 2021 kode criteria (cj) weight (wj) c1 quantity 0.05 c2 integrity 0.20 c3 dedication 0.10 c4 realibility 0.03 c5 initiative 0.09 c6 diligence 0.15 c7 attitude 0.17 c8 motivation 0.08 c9 presence 0.13 summary 1.00 b. implementation this study used a sample of 21 pkh social workers in x sub-district, with the following steps: 1. value tabulation for each of the alternative criteria obtained from interviews with informants and supporting data is shown in table iv. tabel iv. worker data and value tabulation alternative c1 c2 c3 c4 c5 c6 c7 c8 c9 worker 1 a a a a a b+ a b+ a worker 2 a c+ b b+ b b+ b+ b b worker 3 b+ b b b b+ c+ c b c worker 4 b b+ b+ b b+ b b+ b b+ worker 5 b b+ c c b c+ b+ b c+ worker 6 b c+ b c+ b b b b b worker 7 a c+ b b a a b+ a b+ worker 8 a b+ a a b a b+ a a worker 9 a b+ b+ b+ b b a b b worker 10 b a c a d b b b b worker 11 a b a b+ b a b b+ a worker 12 a b b+ b+ b b+ b a b worker 13 a b a a b b b+ b+ b+ worker 14 b d d c d d c d d worker 15 c+ b+ b c+ b c b c+ c+ worker 16 b a a b b b b b a worker 17 a a a a a a a b+ a worker 18 a b b b b b b b b worker 19 c+ b+ b b c+ b b c+ c+ worker 20 b c b b b b+ c+ b b+ worker 21 b b c+ b b+ c+ b+ b c+ 2. the data value for each of the alternative criteria is then converted according to the likert scale of 9-1. so that you get the following results: tabel v. criteria value conversion table alternative c1 c2 c3 c4 c5 c6 c7 c8 c9 worker 1 9 9 9 9 9 7 9 7 9 worker 2 9 4 6 7 6 7 7 6 6 worker 3 7 6 6 6 7 4 3 6 3 worker 4 6 7 7 6 7 6 7 6 7 worker 5 6 7 3 3 6 4 7 6 4 worker 6 6 4 6 4 6 6 6 6 6 worker 7 9 4 6 6 9 9 7 9 7 worker 8 9 7 9 9 6 9 7 9 9 worker 9 9 7 7 7 6 6 9 6 6 worker 10 6 9 3 9 1 6 6 6 6 worker 11 9 6 9 7 6 9 6 7 9 worker 12 9 6 7 7 6 7 6 9 6 worker 13 9 6 9 9 6 6 7 7 7 worker 14 6 1 1 3 1 1 3 1 1 worker 15 4 7 6 4 6 3 6 4 4 worker 16 6 9 9 6 6 6 6 6 9 worker 17 9 9 9 9 9 9 9 7 9 worker 18 9 6 6 6 6 6 6 6 6 worker 19 4 7 6 6 4 6 6 4 4 worker 20 6 3 6 6 6 7 4 6 7 worker 21 6 6 4 6 7 4 7 6 4 3. normalizing the decision matrix using equation (1). if a criterion is included in the profit criteria type, the greater the value the better. meanwhile, if the criteria are included in the type of cost criteria, the smaller the value the better. table vi. types of criteria kode criteria (cj) type c1 quantity benefit c2 integrity benefit c3 dedication benefit c4 realibility benefit c5 initiative benefit c6 diligence benefit c7 attitude benefit c8 motivation benefit c9 presence benefit 1) quantity r11 = = r21 = = 1 … r211 = = 2) integrity r12 = = r22 = = … r212 = = 3) dedication r13 = = r23 = = … r213 = = … 4 | vol.2 no.1, january 2021 9) presence r19 = = r29 = = … r219 = = table vii. normalization results alternative r1 r2 r3 … r9 worker 1 1.00 1.00 1.00 … 1.00 worker 2 1.00 0.44 0.67 … 0.67 worker 3 0.78 0.67 0.67 … 0.33 worker 4 0.67 0.78 0.78 … 0.78 worker 5 0.67 0.78 0.33 … 0.44 worker 6 0.67 0.44 0.67 … 0.67 worker 7 1.00 0.44 0.67 … 0.78 worker 8 1.00 0.78 1.00 … 1.00 worker 9 1.00 0.78 0.78 … 0.67 worker 10 0.67 1.00 0.33 … 0.67 worker 11 1.00 0.67 1.00 … 1.00 worker 12 1.00 0.67 0.78 … 0.67 worker 13 1.00 0.67 1.00 … 0.78 worker 14 0.67 0.11 0.11 … 0.11 worker 15 0.44 0.78 0.67 … 0.44 worker 16 0.67 1.00 1.00 … 1.00 worker 17 1.00 1.00 1.00 … 1.00 worker 18 1.00 0.67 0.67 … 0.67 worker 19 0.44 0.78 0.67 … 0.44 worker 20 0.67 0.33 0.67 … 0.78 worker 21 0.67 0.67 0.44 … 0.44 4. calculating the preference value of each worker using equation (2). vworker1 = (0.05*1.00) + (0.20*1.00) + (0.10*1.00) + (0.03*1.00) + (0.09*1.00) + (0.15*0.78) + (0.17*1.00) + (0.08*0.78) + (0.13*1.00) = 0.95 vworker2 = (0.05*1.00) + (0.20*0.44) + (0.10*0.67) + (0.03*0.78) + (0.09*0.67) + (0.15*0.78) + (0.17*0.78) + (0.08*0.67) + (0.13*0.67) = 0.68 vworker3 = (0.05*0.78) + (0.20*0.67) + (0.10*0.67) + (0.03*0.67) + (0.09*0.78) + (0.15*0.44) + (0.17*0.33) + (0.08*0.67) + (0.13*0.33) = 0.55 vworker4 = (0.05*0.67) + (0.20*0.78) + (0.10*0.78) + (0.03*0.67) + (0.09*0.78) + (0.15*0.67) + (0.17*0.78) + (0.08*0.67) + (0.13*0.78) = 0.74 vworker5 = (0.05*0.67) + (0.20*0.78) + (0.10*0.33) + (0.03*0.33) + (0.09*0.67) + (0.15*0.44) + (0.17*0.78) + (0.08*0.67) + (0.13*0.44) = 0.60 vworker6 = (0.05*0.67) + (0.20*0.44) + (0.10*0.67) + (0.03*0.44) + (0.09*0.67) + (0.15*0.67) + (0.17*0.67) + (0.08*0.67) + (0.13*0.67) = 0.61 vworker7 = (0.05*1.00) + (0.20*0.44) + (0.10*0.67) + (0.03*0.67) + (0.09*1.00) + (0.15*1.00) + (0.17*0.78) + (0.08*1.00) + (0.13*0.78) = 0.78 vworker8 = (0.05*1.00) + (0.20*0.78) + (0.10*1.00) + (0.03*1.00) + (0.09*0.67) + (0.15*1.00) + (0.17*0.78) + (0.08*1.00) + (0.13*1.00) = 0.89 vworker9 = (0.05*1.00) + (0.20*0.78) + (0.10*0.78) + (0.03*0.78) + (0.09*0.67) + (0.15*0.67) + (0.17*1.00) + (0.08*0.67) + (0.13*0.67) = 0.78 vworker10 = (0.05*0.67) + (0.20*1.00) + (0.10*0.33) + (0.03*1.00) + (0.09*0.11) + (0.15*0.67) + (0.17*0.67) + (0.08*0.67) + (0.13*0.67) = 0.65 vworker11 = (0.05*1.00) + (0.20*0.67) + (0.10*1.00) + (0.03*0.78) + (0.09*0.67) + (0.15*1.00) + (0.17*0.67) + (0.08*0.78) + (0.13*1.00) = 0.81 vworker12 = (0.05*1.00) + (0.20*0.67) + (0.10*0.78) + (0.03*0.78) + (0.09*0.67) + (0.15*0.78) + (0.17*0.67) + (0.08*1.00) + (0.13*0.67) = 0.74 vworker13 = (0.05*1.00) + (0.20*0.67) + (0.10*1.00) + (0.03*1.00) + (0.09*0.67) + (0.15*0.67) + (0.17*0.78) + (0.08*0.78) + (0.13*0.78) = 0.76 vworker14 = (0.05*0.67) + (0.20*0.11) + (0.10*0.11) + (0.03*0.33) + (0.09*0.11) + (0.15*0.11) + (0.17*0.33) + (0.08*0.11) + (0.13*0.11) = 0.18 vworker15 = (0.05*0.44) + (0.20*0.78) + (0.10*0.67) + (0.03*0.44) + (0.09*0.67) + (0.15*0.33) + (0.17*0.67) + (0.08*0.44) + (0.13*0.44) = 0.58 vworker16 = (0.05*0.67) + (0.20*1.00) + (0.10*1.00) + (0.03*0.67) + (0.09*0.67) + (0.15*0.67) + (0.17*0.67) + (0.08*0.67) + (0.13*1.00) = 0.80 vworker17 = (0.05*1.00) + (0.20*1.00) + (0.10*1.00) + (0.03*1.00) + (0.09*1.00) + (0.15*1.00) + (0.17*1.00) + (0.08*0.78) + (0.13*1.00) = 0.98 vworker18 = (0.05*1.00) + (0.20*0.67) + (0.10*0.67) + (0.03*0.67) + (0.09*0.67) + (0.15*0.67) + (0.17*0.67) + (0.08*0.67) + (0.13*0.67) = 0.68 vworker19 = (0.05*0.44) + (0.20*0.78) + (0.10*0.67) + (0.03*0.67) + (0.09*0.44) + (0.15*0.67) + (0.17*0.67) + (0.08*0.44) + (0.13*0.44) = 0.62 vworker20 = (0.05*0.67) + (0.20*0.33) + (0.10*0.67) + (0.03*0.67) + (0.09*0.67) + (0.15*0.78) + (0.17*0.44) + (0.08*0.67) + (0.13*0.78) = 0.59 5 | vol.2 no.1, january 2021 vworker21 = (0.05*0.67) + (0.20*0.67) + (0.10*0.44) + (0.03*0.67) + (0.09*0.78) + (0.15*0.44) + (0.17*0.78) + (0.08*0.67) + (0.13*0.44) = 0.60 table viii. preference value of each worker alternative v1 v2 v3 … v9 vi worker 1 0.05 0.20 0.10 … 0.13 0.95 worker 2 0.05 0.09 0.07 … 0.09 0.68 worker 3 0.04 0.13 0.07 … 0.04 0.55 worker 4 0.03 0.16 0.08 … 0.10 0.74 worker 5 0.03 0.16 0.03 … 0.06 0.60 worker 6 0.03 0.09 0.07 … 0.09 0.61 worker 7 0.05 0.09 0.07 … 0.10 0.78 worker 8 0.05 0.16 0.10 … 0.13 0.89 worker 9 0.05 0.16 0.08 … 0.09 0.78 worker 10 0.03 0.20 0.03 … 0.09 0.65 worker 11 0.05 0.13 0.10 … 0.13 0.81 worker 12 0.05 0.13 0.08 … 0.09 0.74 worker 13 0.05 0.13 0.10 … 0.10 0.76 worker 14 0.03 0.02 0.01 … 0.01 0.18 worker 15 0.02 0.16 0.07 … 0.06 0.58 worker 16 0.03 0.20 0.10 … 0.13 0.80 worker 17 0.05 0.20 0.10 … 0.13 0.98 worker 18 0.05 0.13 0.07 … 0.09 0.68 worker 19 0.02 0.16 0.07 … 0.06 0.62 worker 20 0.03 0.07 0.07 … 0.10 0.59 worker 21 0.03 0.13 0.04 … 0.06 0.60 the vi with the largest preference value is the chosen worker, so that worker 17 is the recommended worker to become an exemplary pkh social worker in ppkh x subdistrict. conclusion based on the results described, it is concluded that the implementation of the saw method is effective as a decision support system in determining exemplary pkh social workers in ppkh garut regency. the results of research conducted on 21 pkh social workers in x subdistrict showed that exemplary pkh social worker in the region received a preference value of 0.98. references [1] a. m. siregar, s. faisal, y. cahyana, and b. priyatna, “perbandingan algoritme klasifikasi untuk prediksi cuaca,” j. account. inf. syst., vol. 3, no. 1, pp. 15–24, 2020. [2] f. i. manek, s. faisal, and b. priyatna, “penerapan kmeans clustering untuk mengelompokkan pelanggan berdasarkan data penjualan ayam,” techno xplore j. ilmu komput. dan teknol. inf., vol. 3, no. 2, pp. 88–93, 2018. [3] kementerian sosial ri. 2019. pedoman pelaksanaan program keluarga harapan tahun 2019. direktorat jaminan sosial keluarga. direktorat jenderal perlindungan dan jaminan sosial. [4] penta, mega fidia, dkk. 2019. sistem pendukung keputusan pemilihan karyawan terbaik menggunakan metode saw pada pt. kujang sakti anugrah. jsai (journal scientific and applied informatics), 2(3), 185-192. [5] malau, yesni, dkk. 2018. sistem pendukung keputusan pemilihan pegawai berprestasi di komisi pemilihan umum kabupaten bogor. jurnal teknik komputer, 4(1), 66-73. [6] diana. 2018. metode & aplikasi sistem pendukung keputusan. yogyakarta: deepublish. [7] budiaji, weksi. 2013. skala pengukuran dan jumlah respon skala likert. jipp (jurnal ilmu pertanian dan perikanan), 2(2), 127-133. p-issn : 2715-2448 | e-isssn : 2715-7199 vol.2 no.2 july 2021 buana information tchnology and computer sciences (bit and cs) 31 | vol.2 no.2, july 2021 performance evaluation of adaptive neuro-fuzzy inference system (anfis) in predicting new students (case study : ubp karawang) tatang rohana 1 study program technical information faculty of engineering and computer science, buana perjuangan university email: tatang.rphana@ubpkarawang.ac.id bayu priyatna 2 study program information system faculty of engineering and computer science, buana perjuangan university email: bayu.priyatna@gmail.com ‹β› abstract—the process of admitting new students is an annual routine activity that occurs in a university. this activity is the starting point of the process of searching for prospective new students who meet the criteria expected by the college. one of the colleges that holds new student admissions every year is buana perjuangan university, karawang. there have been several studies that have been conducted on predictions of new students by other researchers, but the results have not been very satisfying, especially problems with the level of accuracy and error. research on anfis studies to predict new students as a solution to the problem of accuracy. this study uses two anfis models, namely backpropagation and hybrid techniques. the application of the adaptive neuro-fuzzy inference system (anfis) model in the predictions of new students at buana perjuangan university, karawang was successful. based on the results of training, the backpropagation technique has an error rate of 0.0394 and the hybrid technique has an error rate of 0.0662. based on the predictive accuracy value that has been done, the backpropagation technique has an accuracy of 4.8 for the value of mean absolute deviation (mad) and 0.156364623 for the value of mean absolute percentage error (mape). meanwhile, based on the mean absolute deviation (mad) value, the backpropagation technique has a value of 0.5 and 0.09516671 for the mean absolute percentage error (mape) value. so it can be concluded that the hybrid technique has a better level of accuracy than the backpropation technique in predicting the number of new students at the university of buana perjuangan karawang keywords— anfis, backpropagation, hybrid, prediction abstract—proses penerimaan mahasiswa baru merupakan kegiatan rutin tahunan yang terjadi di sebuah universitas. kegiatan ini merupakan titik awal dari proses pencarian calon mahasiswa baru yang memenuhi kriteria yang diharapkan oleh perguruan tinggi. salah satu perguruan tinggi yang menyelenggarakan penerimaan mahasiswa baru setiap tahunnya adalah universitas buana perjuangan karawang. ada beberapa penelitian yang telah dilakukan terhadap prediksi mahasiswa baru oleh peneliti lain, namun hasilnya belum terlalu memuaskan, terutama masalah tingkat akurasi dan kesalahan. penelitian tentang studi anfis untuk memprediksi siswa baru sebagai solusi dari masalah akurasi. penelitian ini menggunakan dua model anfis, yaitu teknik backpropagation dan hybrid. penerapan model adaptive neuro-fuzzy inference system (anfis) pada prediksi mahasiswa baru universitas buana perjuangan karawang berhasil. berdasarkan hasil pelatihan, teknik backpropagation memiliki tingkat kesalahan 0,0394 dan teknik hybrid memiliki tingkat kesalahan 0,0662. berdasarkan nilai akurasi prediksi yang telah dilakukan, teknik backpropagation memiliki akurasi sebesar 4,8 untuk nilai mean absolute deviation (mad) dan 0,156364623 untuk nilai mean absolute percentage error (mape). sedangkan berdasarkan nilai mean absolute deviation (mad), teknik backpropagation memiliki nilai 0,5 dan 0,09516671 untuk nilai mean absolute percentage error (mape). sehingga dapat disimpulkan bahwa teknik hybrid memiliki tingkat akurasi yang lebih baik dibandingkan dengan teknik backpropation dalam memprediksi jumlah mahasiswa baru di universitas buana perjuangan karawang. kata kunci— anfis, backpropagation, hybrid, prediksi i. introduction the process of admitting new students is an annual routine activity that occurs in a university. this activity is the starting point of the process of searching for prospective new students who meet the criteria expected by the college. one of the colleges that hold new student admissions every year is buana perjuangan university karawang. buana perjuangan karawang university is one of the universities in the karawang area which is developing very rapidly. this is proven by the high interest of new students who register and can be accepted at buana perjuangan university, karawang. this is of course a challenge and a good opportunity for the university. on the other hand, the stability and availability of campus facilities and infrastructure are things that need to be considered by university administrators. the university certainly has to be able to calculate how many new students are accepted at buana perjuangan university, this is important for the organizers as a material for decision making, especially those related to campus development, infrastructure, and resources that support the teaching and learning process in the campus environment. in this study, the authors used the anfis model to predict the number of new students at the university of buana perjuangan karawang. many studies have been conducted, using the adaptive neuro-fuzzy inference system (anfis) model in the prediction system. among them, the use of artificial neuro fuzzy inference system (anfis) in determining the status of mount merapi activities [5]; the use of backpropagation neural networks for new student 32 | vol.2 no.2, july 2021 admissions in the computer engineering department at sriwijaya state polytechnic [1]; development of artificial neural network model to predict the number of new students in pts surabaya [2]; adaptive neuro fuzzy inference system (anfis) method for prediction of road service levels [11], and other studies. it is expected that the results of this study can provide good accuracy and error rates. in this study, the authors raised the title "adaptive neuro-fuzzy inference system (anfis) study in predicting new student admissions at buana perjuangan university, karawang. ii. method a. types of research in this research, the type of research used is quantitative. the objective of quantitative research is to develop and use mathematical models, theories and hypotheses related to natural phenomena. quantitative research is a type of research that basically uses a deductive-inductive approach. this approach departs from a theoretical framework, the ideas of experts, as well as the understanding of researchers based on their experience, then it is developed into problems and their solutions that are proposed to obtain justification (verification) or an assessment in the form of support for empirical data in the field. b. data collection data is a unit of information recorded by media that can be distinguished from other data, can be analyzed and is relevant to certain programs. data collection is a systematic and standard procedure for obtaining the necessary data. to collect research data, the authors used the interview method. to obtain data sources, the authors conducted interviews with the new student admissions committee, which were then validated with the data center (pusdatin) of buana perjuangan university. c. data analysis the data analysis technique used is to divide the data into two, namely training data and testing data. model training uses training data, while model testing uses testing data. the results of the model trial conclusion will be verified by diagnosis on the testing data. data analysis in this study aims to determine how accurate the adaptive neuro fuzzy inferences system (anfis) model at buana perjuangan university, karawang. in predicting the number of students using the singular value decomposition method. the stages of anfis data analysis can be seen in the following figure. fig. 1. anfis analysis and prediction process the data used in this research is secondary data, namely data on enrollments from new students obtained from the new student admissions committee, and active students for each study program obtained from the center for data and information (pusdatin), university of buana perjuangan karawang. d. framework the problem of admitting new students at a university is a routine problem that occurs every new academic year. so it needs good handlers in its implementation. information on the number of new students is important data for all campus members. both leadership, student affairs, academics, infrastructure and others. this is important because the campus must prepare everything due to the teaching and learning process. the prediction system for new students is certainly very helpful for the campus in preparing for needs a need that must be met by all members of the community in a university. the framework of this prediction system research includes: 1. the number of new students from the 2015/2016 academic year to 2019/2020. 2. the data is obtained from the new student admissions committee and is also equipped from the ubp data center. 3. data preprocessing (initial processing), data cleaning to eliminate data errors and data transformation. 4. the process of training and testing data 5. make predictions for new students 6. evaluate the level of accuracy with the mad and mape models 33 | vol.2 no.2, july 2021 fig. 2. framework e. research subject the data source used in this study is data obtained from the buana perjuangan karawang university new student admissions committee which is then validated with the pusdatin section. new student data used as the source of data in this study were taken from new students from the 2015/2016 academic year to the 2019/2020 academic year. the new student data can be seen in detail in table 1. table 1 new students of ubp karawang 2015 to 2019 year pkn pgsd law psik ak man si if far ti 2015/2026 52 152 146 145 155 240 67 187 97 198 2016/2017 54 165 131 150 179 356 56 178 109 243 2017/2018 61 154 142 164 225 536 77 207 143 352 2018/2019 40 139 156 162 190 490 75 231 140 356 2019/2020 42 148 176 250 208 531 74 209 142 324 source : pusdatin the data will then be processed with a preprocessing process (pre-process) by means of data cleaning and normalization, before being used as data analysis in this study. the preprocessing results will then be processed using the adaptive neuro fuzzy inference system (anfis) method as a method for carrying out the prediction process. iii. results and discussion a. preprocessing data the first step taken in this research data, is to pre-process the data on the number of new students accepted at the university of buana perjuangan for the period of the 2015/2016 academic year to 2019/2020. this needs to be done because the data used in the process is not always ideal for processing. sometimes in this data there are various problems that can interfere with the results of the process itself, such as missing values, redundant data, outliers, or data formats that are not in accordance with the system. the pre-processing of this data includes several stages, including data cleaning and data normalization. 1) data cleaning data cleaning is performed to eliminate inefficient and error-containing data. in this process, data is cleaned using the rapidminer application. fig.3. data cleaning process fig.4. results of data cleaning from the data cleaning process above, it shows that there are no errors or missing in the data used in this study. 2) data normalization before input data is entered into the network, the data is transformed into interval data (normalization). these data are normalized so that they are in the range [0,1]. normalization uses the min max formula (han, kamber, and pei, 2012). (1) x = data input xmin = data x minimum xmax = data x maksimum bmax = upper limit of the interval bmin = lower limit of the interval the purpose of normalization is to equalize the range of values for each data so that each data has a proportional role in each process. to facilitate the conversion process, the name of the study program is changed to the x variable. namely x1, x2, .... x10. table 2 data on the number of new students year x1 x2 x3 x4 x5 x6 x7 x8 x9 x10 2015/2026 52 152 146 145 155 240 67 187 97 198 2016/2017 54 165 131 150 179 356 56 178 109 243 2017/2018 61 154 142 164 225 536 77 207 143 352 2018/2019 40 139 156 162 190 490 75 231 140 356 2019/2020 42 148 176 250 208 531 74 209 142 324 become interval data [0, 1] table 3 convert data to min max r x1 x2 x3 x4 x5 x6 x7 x8 x9 x10 1 0,572 0,5 0,333 0 0 0 0,523 0,169 0 0 2 0,666 1 0 0,047 0,24 0,391 0 0 0,26 0,284 3 1 0,576 0,244 0,18 1 1 1 0,547 1 0,974 4 0 0 0,555 0,161 0,5 0,844 0,904 1 0,934 1 5 0,095 0,346 1 1 0,757 0,983 0,857 0,584 0,978 0,797 34 | vol.2 no.2, july 2021 b. hypothesis test 3) data training process data that has been normalized in the form of min max, then used as a source of data for the training process (training) and testing (test data) in the anfis analysis process. for the anfis analysis process, the data is divided into two parts, namely training data and testing data. new student data from 2015 to 2018 is used as training data, while new student data for 2019 is used as testing data. table 4 data training x1 x2 x3 x4 x5 x6 x7 x8 x9 x10 0,572 0,5 0,333 0 0 0 0,523 0,169 0 0 0,666 1 0 0,047 0,24 0,391 0 0 0,26 0,284 1 0,576 0,244 0,18 1 1 1 0,547 1 0,974 0 0 0,555 0,161 0,5 0,844 0,904 1 0,934 1 table 5 data testing x1 x2 x3 x4 x5 x6 x7 x8 x9 x10 0,095 0,346 1 1 0,757 0,983 0,857 0,584 0,978 0,797 the training process (training) with the adaptive neurofuzzy inference system (anfis) model was carried out using the matlab r2010 tool. data analysis for this prediction uses two adaptive neuro-fuzzy inference system (anfis) models, namely backpropagation and hybrid techniques. so that the training and testing process is also based on these two techniques 4) backpropagation technique data training the data in table 4 shows that new student data is used as training data. the training process consists of 20 epochs (iterations) with an error tolerance of 0.001. from the results of data training with the backpropagation technique, the error rate obtained is 0.0394. fig.5. backpropagation technique training process from the backpropagation technique training above, the resulting error rate is 0.0394 with an error tolerance of 0.001 with 20 iterations. from this error value, the rate of increase and decrease in student predictions using the backpropagation technique is 3.94%. 5) hybrid technique data training in the same way, subsequent training is carried out using hybrid techniques. the training process was carried out 20 times with an error tolerance of 0.001. the training process can be seen in the following image: fig. 6. hybrid technique training process based on the results of training with the hybrid technique, an error rate of 0.0662 was generated. with this error value, the predicted value of increase and decrease in new students with the hybrid technique is 6.62%. 6) data testing process the next process is the testing process or data testing. this test is an application of the results of training data to predictions of new students. the data used in this process is the number of new students in 2019, the data can be seen in the following table: table 6 data testing x1 x2 x3 x4 x5 x6 x7 x8 x9 x10 0,095 0,346 1 1 0,757 0,983 0,857 0,584 0,978 0,797 42 148 176 250 208 531 74 209 142 324 from the training process above, the predictive value obtained with the backpropagation technique is 0.0394 while the hybrid technique is 0.0662. 7) backpropagation data testing the test results with the parameters obtained from the data training process with an error rate of 0.0394 for the backpropagation technique, then the prediction results and errors of new students for 2019 are obtained. table 7 backpropagation technique prediction and error result no faculty data aktual prediction error (backpropagation) 1 pkn 42 42 0 2 pgsd 148 144 4 3 law 176 162 14 4 psik 250 168 82 5 ak 208 197 11 6 man 531 509 22 7 si 74 78 4 8 if 209 240 31 9 far 142 146 4 10 ti 324 370 46 fig 7. prediction of backpropagation technique 35 | vol.2 no.2, july 2021 8) hybrid data testing meanwhile, from the data training process with an error rate of 0.0662 for the hybrid technique, the results of predictions and errors for new students for 2019 are as follows: table 8 hybrid technique prediction and error results no faculty data aktual prediction error (hybrid) 1 pkn 42 43 1 2 pgsd 148 148 0 3 law 176 166 10 4 psik 250 173 77 5 ak 208 203 5 6 man 531 522 9 7 si 74 80 6 8 if 209 246 37 9 far 142 149 7 10 ti 324 379 55 fig 8. prediction of hybrid techniques c. analysis and discussion the model used in this study is the adaptive neuro fuzzy inference system (anfis). meanwhile, the techniques used in the fuzzy inference system (fis) are backpropagation and hybrid techniques. to measure the level of accuracy of the two techniques, the error rate of each technique must be sought. (hanke & wichern, 2005) said that forecasting techniques that use quantitative data often contain data in the form of a certain time series. which is where there are errors / errors made by forecasting techniques. therefore a method is needed to measure how much error / error can be generated by forecasting methods to be reconsidered before making a decision. there are also uses of this method of measuring error forecasting are: ▪ comparing the accuracy of the 2 (or more) forecasting methods used. ▪ measuring the reliability and benefits of the forecasting method used. ▪ finding the optimal forecasting method for the organization or company. to measure the level of accuracy of the backpropagation and hybrid techniques, the mean absolute deviation (mad) and mean absolute percentage error (mape) methods are used. a good level of accuracy is if the error rate is smaller than the others. 9) mean absolute deviation (mad) mean absolute deviation measures the accuracy of the prediction (forecast) by making an equal of the magnitude of the forecast error, where each prediction has an absolute value for each error. the formula used to calculate mad is: (2) information : y = actual value in period t y ̈t = forecast value in period t n = number of data periods from the measurement results of the accuracy level of the backpropagation and hybrid techniques that have been carried out using the mean absolute deviation (mad) method, the following values are obtained: table 9 mean absolute deviation hybrid aktual hybrid y1-ŷt 42 42 -1 148 144 0 176 162 10 250 168 77 208 197 5 531 509 9 74 78 -6 209 240 -37 142 146 -7 324 370 -55 total absolute deviation 5 mad 0,5 table 10 mean absolute deviation (mad) backpropagation aktual backpropagation y1-ŷt 42 43 0 148 148 4 176 166 14 250 173 82 208 203 11 531 522 22 74 80 -4 209 246 -31 142 149 -4 324 379 -46 jumlah deviasi absolut 48 mad 4,8 from the table above, backpropagation has a mean absolute deviation (mad) value obtained of 4.8, and for the total deviation of 48. as for the hybrid technique, the mean absolute deviation (mad) value obtained is 0.5 and the number of absolute deviations is 5. 10) mean absolute percentage error (mape) the mean absolute percentage error is calculated by finding the error / absolute error in each period, which is divided by the actual observed value for that period, and an average of the absolute percentage errors is made. the formula used to calculate mape is: (3) information : n = the number of data periods yt = actual value in period t yt = the forecast value in period t based on the results of the calculation of mean absolute percentage error (mape) for the backpropagation technique, the error value is 1.56365 with a mape value of 0.1563647. with a prediction error value of 15.6%. 36 | vol.2 no.2, july 2021 table 11 mean absolute percentage error backpropagation aktual backpropagation y1-ŷt 42 42 0 148 144 0,02703 176 162 0,07955 250 168 0,328 208 197 0,05288 531 509 0,04143 74 78 -0,05405 209 240 -0,14833 142 146 -0,02817 324 370 -0,14198 jumlah deviasi absolut 1,56363 mape 0 as for the hybrid technique, the error value is 0.95167162 with a mape value of 0.09516671. with this, the prediction error value with the hybrid technique is 9.52%. table 12 mean absolute percentage error hybrid aktual hybrid y1-ŷt 42 43 -0,023809 148 148 0 176 166 0,056818 250 173 0,308 208 203 0,024038 531 522 0,016949 74 80 -0,081081 209 246 -0,177033 142 149 -0,0492957 324 379 -0,169753 jumlah deviasi absolut 0,95167162 mape 0,09516671 with the results of this test, it can be concluded that the predictions of new students at the university of buana perjuangan karawang with the adaptive neuro-fuzzy inference system (anfis) model can be used properly. this is evidenced by the results of training (training) backpropagation technique has an error rate (error rate) of 0.0394, while the hybrid technique has an error rate of 0.0662. then based on the calculation of the value of accuracy in predicting, the backpropagation technique has an accuracy of 4.8 for the value of mean absolute deviation (mad) and 0.156364623 for the value of mean absolute percentage error (mape). meanwhile, based on the mean absolute deviation (mad) value, the backpropagation technique has a value of 0.5 and 0.09516671 for the mean absolute percentage error (mape) value. when compared to the accuracy, the hybrid technique is more accurate than the backpropagation technique, because it has a smaller error accuracy value. iii. conclusion from the results of research and testing that have been carried out on the study of adaptive neuro-fuzzy inference system (anfis) in predicting new students at buana perjuangan university, karawang, it can be concluded as follows: 1. the application of the adaptive neuro-fuzzy inference system (anfis) model in the predictions of new students at buana perjuangan university, karawang is successful. based on the results of training, the backpropagation technique has an error rate of 0.0394 and the hybrid technique has an error rate of 0.0662. 2. based on the predictive accuracy value that has been done, the backpropagation technique has an accuracy of 4.8 for the value of mean absolute deviation (mad) and 0.156364623 for the value of mean absolute percentage error (mape). meanwhile, based on the mean absolute deviation (mad) value, the backpropagation technique has a value of 0.5 and 0.09516671 for the mean absolute percentage error (mape) value. 3. based on the accuracy results, the hybrid technique is more accurate than the backpropagation technique, because it has a smaller predictive error accuracy value based on the mean absolute deviation (mad) and mean absolute percentage error (mape) values. acknowledgment finally, the authors would like to thank all those who have helped and provided criticism and suggestions so that this research can be completed on time. references [1] agustin, maria, ”penggunaan jaringan syaraf tiruan backpropagation untuk penerimaan mahasiswa baru pada jurusan teknik komputer di politeknik negeri sriwijaya”, jurnal : jurusan teknik komputer politeknik negeri sriwijaya.2012 [2] aldrian. e, setiawan. j.d. “ application of multivariate anfis for daily rainfall prediction: influences of training data size”. makara, sains, volume 12, no. 1, april 2008: 7-14 [3] alven safik ritonga1, suryo atmojo, “ pengembangan model jaringan syaraf tiruan untuk memprediksi jumlah mahasiswa baru di pts surabaya (studi kasus universitas wijaya putra) “, seminar nasional teknik industri. 2017 [4] a. rahman, a.g. abdullah, dan d.l. hakim, “prakiraan beban puncak jangka panjang pada sistem kelistrikan indonesia menggunakan algoritma adaptive neuro-fuzzy inference sistem”, electrans, vol.11, no.2, 18 -26. 2012, [5] anugrah. ”perbandingan jaringan saraf tiruan backpropagation dan metode deret berkala box-jenkins (arima) sebagai metode peramalan”, jurnal: jurusan matematika fakultas mipa universitas negeri semarang. 2012 [6] bagus, fatkhurrozi, m. aziz muslim, didik r. santoso, “penggunaan artificial neuro fuzzy inference sistem (anfis) dalam penentuan status aktivitas gunung merapi”, jurnal eeccis vol. 6, no. 2. 2012 [7] candra, dewi, werdha, wilubertha, himawati, “prediksi tingkat pengangguran menggunakan adaptif neuro fuzzy inference system (anfis)”, konferensi nasional sistem & informatika. , 2015 [8] f. fanita, z. rustam. “predicting the jakarta composite index price using anfis and classifying prediction result based on relative error by fuzzy kernel c-means. aip conference proceedings 2023, 020206 2018 [9] gandhi, ramadhona, budi, darma, setiawan, fitra a. bachtiar, “ prediksi produktivitas padi menggunkan jaringan syaraf tiruan backpropagation”, jurnal pengembangan teknologi informasi dan ilmu komputer. 2012 [10] hanke, j.e., wichern, d.w., “ business forecasting “. prentice hall, new york. 2015 [11] husein. a. m, simarmata. a. m. “drug demand prediction model using adaptive neuro fuzzy inference system (anfis). journal publications & informatics engineering research, volume 4, number 1, october 2019 [12] linda, puspa “peramalan penjualan produksi teh botol sosro pada pt. sinar sosro sumatera bagian utara tahun 2014 dengan metode arima box-jenkins”, jurnal : fakultas matematika dan ipa universitas sumatera utara. 2015 [13] mat, a, zulkarnaen n.k, deris. s, nurul. n.h, saberi. m.m, “a review on predictive modeling technique for student academic performance monitoring”. matec web of conferences . 2019 37 | vol.2 no.2, july 2021 [14] matonda, arizona, zakso, “jaringan syaraf tiruan dengan algoritma backpropagation untuk penentuan kelulusan sidang skripsi”, jakarta: pelita informatika budi darma. 2013 [15] mulyani. d. “prediction of new student numbers using least square method”. (ijarai) international journal of advanced research in artificial intelligence, vol. 4, no.11, 2015 [16] noor, azizah, kusworo, adi b, achmad, widodo.“ metode adaptive neuro fuzzy inference system (anfis) untuk prediksi tingkat layanan jalan “,jurnal sistem informasi bisnis 03. 2013 [17] risa, helilintar, intan nur farida, “ penerapan algoritma k-means clustering untuk prediksi prestasi nilai akademik mahasiwa “, jurnal sains dan informatika volume 4, nomor 2. 2018 [18] rohana, tatang . “ teknik pengolahan citra dan adaptive neuro fuzzy inference system (anfis) untuk mendeteksi cacat keping printed circuit board (pcb)” jakarta. 2012 [19] siang, jong jek, “jaringan syaraf tiruan dan pemogramannya menggunakan matlab”. yogyakarta: andi offset. 2005 [20] septiawan, f. y., dewa, c. k., afiahayati. “prediction of currency exchange rate in forex trading system using genetic algorithm”. international interdisciplinary conference on science technology engineering management pharmacy and humanities held on 22nd – 23rd. isbn: 9780998900001. 2017 [21] slim. a, hush. d, ojah. d. “predicting student enrollment based on student and college characteristics. proceedings of the 11th international conference on educational data mining [22] tiwari. s, babbar. r , kaur. g. ”performance evaluation of two anfis models for predicting water quality index of river satluj (india)”. advances in civil engineering, volume 2018 [23] wanti, rahayu1, “ model penentuan guru berprestasi berbasis adaptive neuro fuzzy inference system (anfis), jurnal sisfotek global 2088 – 1762 vol. 7 no. 1. 2017 [24] wiwik, anggraeni, “aplikasi jaringan syaraf tiruan untuk peramalan permintaan barang”, jurnal: jurusan sistem informasi, institut teknologi sepuluh november. 2012 [25] widodo, p.p, handayanto, r.t., “ penerapan soft computing dengan matlab “, bandung. 2012, p-issn : 2715-2448 | e-issn : 2715-7199 vol.2 no.2, july 2021 buana information tchnology and computer sciences (bit and cs) 62 | vol.2 no.2, july 2021 the use of usability tests in website-based student report value processing information systems arip solehudin 1 technical information faculty of computer science, university singaperbangsa karawang, indonesia arip.solehudin@staff.unsika.ac.id nono heryana 2 information system faculty of computer science, university singaperbangsa karawang, indonesia nono@unsika.ac.id ‹β› rieke retnosary 3 information system faculty of computer science, university buana perjuangan karawang, indonesia riekeretnosary@ubpkarawang.ac.id abstract the information system for processing student report cards based on the website is a system that provides information on student activity reports online in the form of grade reports and related student information based on the web, thus helping speed and quality in delivering information. problems that occur in processing report cards at smpn 2 telagasari are currently still using report cards and inputting student scores is still manual. this study aims to build a value information system that makes it easier to check, record, and report computerized student grade data. in addition, by being web-based, data information can be accessed at any time. this website uses xampp as a web server for system design and mysql as a database. the design of the login menu consisting of admins, teachers, and students has separate access when opening the application so that the security of the program is maintained. testing this website using black-box testing and likert scale usability testing to make it easier for authors to assess whether this online report card processing information system website can facilitate users. keywords: information system, value, report card, website, mysql abstrak sistem informasi pengolahan raport siswa berbasis website merupakan sistem yang menyediakan informasi laporan kegiatan siswa secara online berupa raport nilai dan informasi siswa terkait berbasis web, sehingga membantu kecepatan dan kualitas dalam penyampaian informasi. permasalahan yang terjadi dalam pengolahan raport di smpn 2 telagasari saat ini masih menggunakan raport dan penginputan nilai siswa masih manual. penelitian ini bertujuan untuk membangun sistem informasi nilai yang memudahkan pengecekan, pencatatan, dan pelaporan data nilai siswa secara komputerisasi. selain itu, dengan berbasis web, informasi data dapat diakses setiap saat. website ini menggunakan xampp sebagai web server untuk perancangan sistem dan mysql sebagai database. perancangan menu login yang terdiri dari admin, guru, dan siswa memiliki akses tersendiri saat membuka aplikasi sehingga keamanan program tetap terjaga. pengujian website ini menggunakan pengujian black box dan pengujian usability skala likert untuk memudahkan penulis menilai apakah website sistem informasi pengolahan raport online ini dapat memudahkan pengguna. kata kunci: sistem informasi, nilai, rapor, website, mysql i. preliminary the use of technology and information is now also an aspect of life that is not only limited to the work environment. this is what makes technology and data very important for the survival of society. school is an environment that has used information technology. education is an institution that aims to provide services in the academic aspect that is used by many people in guiding a student. the school directs students to become people who can advance the nation. the school is one of the bodies devoted to guiding students in the monitoring of educators. at the end of each semester, the school evaluates a student in one semester as a basis for measuring academic progress that has been passed. each teacher will process these values and then will be given to each homeroom teacher respectively. each homeroom teacher then collects and produces in one form an evaluation document, which results in a document known as a student report card. the processing of report cards at smp negeri 2 telagasari is arguably less efficient and not optimal. each subject teacher gives grades to students which are then handed over to the homeroom teacher to be processed into a diploma (report). in the process of ensuring the student report card numbers continue to use the old method. subject teachers share student grades with homeroom teachers using excel separately. this makes it difficult for the homeroom teacher to rework the information submitted by the teacher on the subject, as a result, the way the report card works is held back for each grade. the software used does not match the wishes of consumers as a result, it causes consumers to have difficulty in working on student report cards. this of course causes the process of processing information and making information, from the processing field not so optimal, it can limit the way of searching and presenting the required information. to improve the quality of learning, increase the effectiveness of the duration and source of energy for schools, both in 63 | vol.2 no.2, july 2021 guiding practice activities or school administration, such as in working on report cards, the use of information technology is highly desirable. in collecting, calculating student grades, and printing student report cards, this report card processing information system will greatly facilitate the process of making student report cards. ii. research methods a. information systems an information system is a system that is produced by people who are divided into parts into a body to achieve the goal of presenting data [1]. b. data processing an information system is a system that is produced by people who are divided into parts into a body to achieve the goal of presenting data [2]. c. website website or short for the web can be known as a group of pages that are divided into several pages that contain data in the form of virtual information in the form of text, paintings, films, audio, and other animations that are carried out via connection routes [3]. d. report card report cards are information on the results of teaching participants' information while exploring teaching and learning activities and submitted at the end of the activity twice a year whose reporting, in this case, is the results of daily quizzes, daily obligations, midterm tests, end-of-semester tests, character, extracurricular along with the required records related to reporting cards [4]. e. student students or students are those who are specially handed over by their parents to explore training activities held at the school, which aim to become a creature who has insight, skills, professionalism, character, has high morals, and is independent [5]. f. teacher the teacher or teacher is a person who guides and provides teaching because of his rights and obligations he is responsible for the learning of teaching participants [6]. g. school administration school administration is a sub-system of the body, in this case, it is a school body. the key activity is to take care of all forms of school administration, from correspondence to record keeping. so the rules of effort are not only related to correspondence activities but also relate to all explanatory materials and data in the form of scripts [7]. h. anfield modeling language (uml) uml (unified modeling language) is a description that is widely used in industry to describe requirements, make system analysis and design, and describe systems in objectoriented programming [8]. the following are the types of uml: 1. use case diagram use case diagram is a model for data system actions to be made. use case charts can be used to identify the uses of everything contained in the data system and who has the power to use those uses [9]. 2. class diagram the class chart describes the state of the system usability and requirements related to important menus and database connections [10]. 3. activity diagram activity diagrams describe a workflow (activity movement) or activity (activity) of a system or business field [9]. 4. sequence diagram sequence diagrams describe the subject's actions against a use case by defining the duration of the subject's life and the notes to be sent and obtained by fellow-subjects [9]. i. mysql mysql is a relational database management system. that is information that is managed through a database that will be placed in several separate tables so that dealing with information will be much faster. mysql can also be used to manage databases from the smallest to the largest[11]. j. database a database is a collection of interrelated information, hidden in external funds and used as special software to manipulate it. the database is also a significant part of the information system because it acts as a data facilitator for its users [12]. k. waterfall the modified waterfall form is a sequential concept method, often used in software development methods, in its progress, it always flows to the bottom (like a waterfall) through the levels of communication, planning, modeling, construction, and deployment [13]. l. php php is a scripting programming language that was originally developed to generate html statements. moreover, programming developed with php one hundred percent is always shown in the form of html code [14]. m. codeigniter code igniter is a framework from php that is opensource using the mvc (visual meteorological condition) procedure to make it easier for developers or programmers to create an application with a website platform without having to make it from scratch[1]. n. black box testing black box testing is centered on a functional element of the software. the tester can interpret the combination of an input situation and carry out tests on elements of a functional program. black box testing is not an alternative solution to white box testing, however, an accessory for testing situations that are not covered by white box testing[15]. iii. results and discussion to get an analysis of user needs for the system, likert scale usability is needed as a method to be able to analyze user needs in using this system, the following steps are: 1. the likert scale usability test focuses on the convenience of consumers in using the student report card data system, as a result, the instrument used for this research is usability research. in the usability aspect, the test uses an evaluation sheet in the form of a questionnaire or questionnaire which will be distributed to respondents directly after trying the information system. the questionnaire used is the use questionnaire by lund a.m (2001) which already has four criteria, namely usefulness, ease of learning, and satisfaction. in the calculation process, 64 | vol.2 no.2, july 2021 the questionnaire has five scales that are used as benchmarks, including strongly agree (ss), agree (s), uncertain (rg), disagree (ts), and strongly disagree (sts). the following is a table of likert scale usability questions: table 1 questionnaire likert scale 2. data analysis techniques are used to determine the level of achievement of the objectives of the research based on the data that has been collected. testing this usability aspect by using likert scale quantitative data analysis. the likert scale contained in the use questionnaire instrument can use answers on a scale of five or scale seven, in this study using a scale of five. according to sugiyono (2009), the answers to each instrument using a likert scale have a gradation from very positive to negative. the value of 1 is the smallest while the value of 5 is the largest. table 2 likert skala scale classification the total value obtained is then calculated by the following formula. 𝑝𝑒𝑟𝑠𝑒𝑛𝑡𝑎𝑠𝑒 (%) = total value maximal value information : total value = total value obtained from respondents' answers maximum value = number of statements x number of respondents x 5 after getting the results of the calculations, the values obtained will then be converted into qualitative values in the percentage assessment table. before knowing the percentage assessment table first, look for the interval distance rating of the likert scale using the following formula. interval = 100/total score (likert) = 100/5 = 20 from the calculation of the interval above, it can be seen that the result of the distance interval for the assessment percentage table is 20, so the percentage assessment table can be seen in table 3. table 3 rating percentage 3. furthermore, the results of testing data from the four tables are collected and accumulated so that it will produce a summary of the usability testing of the student report card processing information system with a total score that can be seen in the following table. table 4 results of usability test accumulated data 𝑥100% 65 | vol.2 no.2, july 2021 based on the results of the accumulated usability test data in table 4.33, a total score of 2538 points was obtained from a maximum possible score of 3000 points. from the results obtained, the percentage of eligibility is calculated based on the data obtained. the calculation of the percentage of eligibility based on the data is as follows. (%) = 𝑇𝑜𝑡𝑎𝑙 𝑆𝑐𝑜𝑟𝑒 𝑀𝑎𝑥𝑖𝑚𝑎𝑙 𝑆𝑐𝑜𝑟𝑒𝑙 × 100% = 2538 3000 × 100% = 84,6 % the result of calculating the percentage of eligibility is 84.6% so it can be concluded that the student report card processing information system meets the usability standard with the "very good" category when viewed in the feasibility percentage table contained in chapter 3. a usability test is a test used in beta testing is a stage last in the construction process. with input and suggestions, the student report card processing information system will continue to be developed based on user evaluations so that this student report card processing information system can achieve the maximum level of feasibility. after getting the results of the likert scale usability testing related to data requirements to determine the level of user interest, the next stage is the development and design of the system following the stages in making the system: 1. in designing this system using a use case diagram as a modeling system, which was created using uml (unified modeling language). describe in building a system program. it is used as a medium that has a function to design the system and describe the interaction between the actor and the system. the following is an example of a system diagram use case design: figure 1 use case diagram 2. activity diagram of system capital is useful in describing the workflow on the system such as describing login activities, data management activities, modifying data, and deleting data. the following is an example of an activity diagram that has been created. figure 2 activity diagram 3. squance diagrams are useful for compiling the steps of messages sent to find out the flow of the relationship between objects because the sequence is the most useful diagram for dividing the use case model into a barangin system specification. the following squance example diagram for designing a goods lending system can be seen below. figure 3 sequence diagram 4. a class diagram is a modeling of the system design structure contained in the database. class diagram also aims to describe the functions needed in designing and creating systems. the following is an example of a class diagram for designing a goods lending system, which can be seen below: 66 | vol.2 no.2, july 2021 figure 4 class diagram 5. system interface design in designing this user interface, it is designed using a pencil application. user interface design aims to describe the appearance of the system to be created. the following is the user interface design of the goods asset inventory system: login interface design figure 5 login interface design figure 6 admin interface design figure 7 teacher interface design figure 7 student interface design iv. conclusion based on the results of research and discussion, the following conclusions can be drawn: 1. 1. the student report card processing information system uses the waterfall development model from pressman with the stages of communication, planning, modeling, construction, and deployment. the information system developed has features to manage teachers, manage students, manage classes, manage subjects, manage grades, and print student report cards. 2. the developed information system has carried out tests that focus on usability aspects with the calculation results in the very good category, the product is developed to facilitate the user's work so that the objectives of the research are achieved. reference [1] m. destiningrum and q. j. adrian, “sistem informasi penjadwalan dokter berbassis web dengan menggunakan framework codeigniter (studi kasus: rumah sakit yukum medical centre),” j. teknoinfo, vol. 11, no. 2, p. 30, 2017, doi: 10.33365/jti.v11i2.24. [2] m. h. a. muhdar abdurahman1, mudar safi2, “ijis indonesian journal on information system issn 25486438,” ijis-indonesia j. inf. syst., vol. 4, no. april, pp. 69–76, 2019, [online]. available: https://media.neliti.com/media/publications/260171sistem-informasi-pengolahan-data-pembelie5ea5a2b.pdf. [3] a. josi, “penerapan metode prototyping dalam membangun website desa (studi kasus desa sugihan kecamatan rambang),” jti, vol. 9, no. 1, pp. 50–57, 2017. [4] a. helmy syahrizal, “strategi optimalisasi pengelolaan kekayaan (aset) desa dalam pembangunan desa (studi kasus di desa sambiroto kecamatan kapas kabupaten bojonegoro),” publika, vol. 6, no. 4, 2018. [5] e. r. dan e. yuliawati, “pengembangan produk lampu meja belajar dengan metode kano dan quality function deployment (qfd),” j. res. technol., vol. 2, no. 2, pp. 78–86, 2016. [6] n. m. janna, “konsep uji validitas dan reliabilitas dengan menggunakan spss,” artik. sekol. tinggi agama islam darul dakwah wal-irsyad kota makassar, no. 18210047, pp. 1–13, 2020. [7] e. santika, “pengertian dan proses administrasi ketatausahaan sekolah,” no. 18029105, pp. 1–5, 2020, 67 | vol.2 no.2, july 2021 doi: 10.31227/osf.io/3q4xm. [8] d. susianto, “perancangan sistem pemesanan e-tiket pada wisata di lampung berbasis web mobil,” vol. 2, pp. 60–71, 2019. [9] k. kawano, y. umemura, and y. kano, " field assessment and inheritance of cassava resistance to superelongation disease 1," crop sci., vol. 23, no. 2, pp. 201–205, 1983, doi: 10.2135/cropsci1983.0011183x002300020002x. [10] b. huda, “sistem informasi data penduduk berbasis android dan web monitoring studi kasus pemerintah kota karawang (penelitian dilakukan di kab. karawang),” buana ilmu, vol. 3, no. 1, pp. 62–69, 2018, doi: 10.36805/bi.v3i1.456. [11] d. sanjaya, h. abdurachman, a. a. wicaksono, and f. masya, “sistem informasi pengendalian asset kendaraan di perusahaan transportasi,” rabit j. teknol. dan sist. inf. univrab, vol. 6, no. 1, pp. 24–32, 2021, doi: 10.36341/rabit.v6i1.1544. [12] a. andaru, “pengertian database secara umum,” osf prepr., p. 2, 2018. [13] h. kurniawan, w. apriliah, i. kurnia, and d. firmansyah, “penerapan metode waterfall dalam perancangan sistem informasi penggajian pada smk bina karya karawang,” j. interkom j. publ. ilm. bid. teknol. inf. dan komun., vol. 14, no. 4, pp. 13–23, 2021, doi: 10.35969/interkom.v14i4.78. [14] n. n. abdur rochman,achmad sidik, “perancangan sistem informasi administrasi pembayaran spp siswa berbasis web di smk al-amanah,” j. sisfotek glob., vol. 8, no. 1, pp. 52–52, 2018. [15] q. mardzotillah and m. ridwan, “sistem tracer study dan persebaran alumni berbasis web di universitas islam syekh-yusuf tangerang,” jutis (jurnal tek. inform., vol. 8, no. 1, pp. 90–106, 2020, [online]. available: http://ejournal.unis.ac.id/index.php/jutis/article/view/70 5. p-issn : 2715-2448 | e-isssn : 2715-7199 vol.2 no.2 july 2021 buana information tchnology and computer sciences (bit and cs) 55 | vol.2 no.2, july 2021 implementation location-based service (lbs) on mobile application for searching dormitory moh hasan basri 1 information system faculty of engineering and computer science university buana perjuangan karawang si17.mohbasri@mhs.ubpkarawang.ac.id aprilia hananto 2 information system faculty of engineering and computer science university buana perjuangan karawang aprilia.hananto@ubpkarawang.ac.id ‹β› siti masruroh 3 information system faculty of engineering and computer science university buana perjuangan karawang siti.masruroh@ubpkarawang.ac.id abstract—the growth of mobile features, especially smartphones at this time can be said to be very developed. judging from the increasing number of mobile users, the availability of various mobile features is also growing very rapidly. in line with these facts, the author aims to create a mobile dormitory application using location-based service (lbs) technology with the application of the haversine method, making it easier for users to find the closest dormitoryto their place of work with appropriate facilities. this application is not only to search for dormitorys, it can also place orders online, this is also an opportunity for the manager or owner of the dormitoryto promote their dormitory. methods of data collection are done by observation, interviews, literature studies, and documentation. the development of this application system uses the waterfall method. the final result of this research will create an android-based dormitory mobile application that can make it easier for dormitoryseekers to search and book dormitorys as well as a means of promotion for dormitoryowners and managers. keywords—android, haversine, location-based service, mobile dormitory, waterfall abstrak—pertumbuhan fitur mobile khususnya smartphone saat ini dapat dikatakan sangat berkembang. dilihat dari peningkatan jumlah pengguna ponsel, ketersediaan berbagai fitur ponsel juga berkembang sangat pesat. sejalan dengan fakta tersebut, penulis bertujuan untuk membuat aplikasi mobile kost menggunakan teknologi location based service (lbs) dengan penerapan metode haversine, sehingga memudahkan pengguna untuk mencari kost terdekat dengan tempat kerjanya dengan tepat. fasilitas. aplikasi ini tidak hanya untuk mencari tempat kost, juga dapat melakukan pemesanan secara online, hal ini juga menjadi peluang bagi pengelola atau pemilik kost untuk mempromosikan kostnya. metode pengumpulan data dilakukan dengan observasi, wawancara, studi pustaka, dan dokumentasi. pembangunan sistem aplikasi ini menggunakan metode waterfall. hasil akhir dari penelitian ini akan membuat sebuah aplikasi mobile kost berbasis android yang dapat mempermudah pencari kost untuk mencari dan memesan kost serta sebagai sarana promosi bagi pemilik dan pengelola kost. kata kunci—android, haversine, layanan berbasis lokasi, mobile dormitory, waterfall i. preliminary karawang regency is an area in west java province which has an area of 1,753.27 km2 or 3.73 percent of the area of west java province with a population of 2,336,009 people [1]. karawang is known as a big industrial city in indonesia. this can be seen from the number of companies that are established in karawang. therefore, many immigrants from various cities come to work for companies in this city. migrants who mostly don't know the area in karawang generally have problems finding dormitorys around the company where they work, especially dormitorys with the closest distance from the company. on the other hand, dormitoryowners find it difficult to promote or publicize available rooms or dormitorys. the growth of mobile features, especially smartphones at this time, can be said to be very developed. judging from the increasing number of mobile users, the availability of mobile features is also growing very rapidly and internet technology has progressed very drastically. the internet has become a very effective means of information and communication. with the internet, various information in the world can be obtained quickly. currently, android applications are widely used in various fields, one of which is in the field of business which has implemented many android applications and has been proven to provide benefits to the community [2]. based on this background, the author aims to design an application with the title "mobile dormitory application karawang using location-based service (lbs)", which can be used as a means to assist users in finding dormitorys based on the closest distance and knowing the dormitory address and information about other dormitorys. , and on the application, you can order dormitorys online. with the presence of this dormitory application, users can search for 56 | vol.2 no.2, july 2021 the nearest dormitoryand book a dormitory. this application is made with the java programming language based on android and to determine the closest distance by utilizing location based service (lbs) using the haversine formula [3]. ii. research methods the methods used in this study are data collection methods, waterfall system development methods, haversine formula to determine the closest distance in finding a dormitoryin the application built, and location-based service to connect the user's location by utilizing the global positioning system which is already available on mobile devices. 4]. 1. method of collecting data in the data collection method, the researchers used four ways, including: 1)observation researchers visited and made direct observations to several dormitorys in karawang, these observations were like seeing and checking the overall state of the dormitorysuch as the state of the dormitory rooms and available facilities. the results of observations are in the form of facts and information about the state of the dormitory. 2)interview conduct direct interviews with dormitoryowners about how to order, pay and market the dormitoryand ask for information about the dormitorysuch as the name of the dormitory, dormitory address, dormitory rules, and others. the results of this interview are in the form of dormitory data that can be used as material for designing applications to be made. then, interview the dormitoryseekers about how or the methods used so far to search for dormitorys and how to order them. the results of interviews with dormitoryresidents are in the form of information that can later be used as the need for designing applications to be made. 3) literature study this literature study was conducted to obtain theoretical references related to the research topic raised, this theory can be obtained from several sources such as journals and theses. 4) documentation the documentation carried out in this study such as taking pictures of the dormitorywith a smartphone camera, existing facilities, and other images that support the process of this research. 2. system development method the development of this system uses the waterfall method [5], while the stages in the development of this system are as follows: figure 1 model waterfall 1) needs analysis (requirements analysis) at this stage, intensive requirements collection is carried out to determine software requirements so that users can understand the type of software needed [6]. 2) design (design) system design is a stage that focuses on the appearance of the system, including data structures, system software architecture, and system interfaces. this stage is designed to meet user needs by using a system in the form of designing a mobile application system display, such as searching for dormitory rooms based on location or location-based service (lbs) with the haversine method [3]. 3) implementation (coding) this stage is the coding stage of the program which is the implementation process in the form of an order or the realization process of the command form, and the computer can use a programming language to understand the process. the mobile application system that will be created uses the java programming language using android studio and the firebase database. this implementation phase includes the application of the haversine formula to determine the closest dormitory distance in the application [7]. 4) testing (testing) this testing stage is to ensure that the system that has been completed is by the designer's design, to find out whether the implemented functions can be used in the process of making and designing the mobile application system. 3. haversine formula the harvesine method is used to calculate the longitude of two points on the earth's surface based on latitude and longitude. haversine formula requires inputting the longitude and latitude of the user's location. the following is the formula of haversine [8]. figure 2 haversine formula information : lat1 = degree latitude of starting point long1 = degree longitude starting point lat2 = degree latitude of destination point long2 = degree longitude of destination point x = longitude (longitude) y = latitude (latitude) d = distance (km) 1 degree = 0.0174532925 radians r = 6371 km from the above formula, the following is a distance calculation using the haversine formula and an analysis of 57 | vol.2 no.2, july 2021 the haversine formula calculation. this calculation is to measure the distance by using a sample of two locations, namely from the starting point of karawang international industrial city to the destination point of griya kost. location 1 karawang international industrial city is known : latitude = -6.359197486564478 longitude = 107.2742426711646 location 2 griya kost is known : latitude = -6.352268604491233 longitude = 107.309655751302 difference in longitude locations 1 and 2 difference = longitude location 1 longitude location 2 difference = 107.2742426711646–107.3096557511302 difference = -0.03541307997 convert latitude location 1,location 2,longitude difference to rad conversion result latitude location 1 = -0.110988933924 rad latitude location 2 = -0.110868002119 rad . conversion result longitude difference conversion result = -0.000618074843 rad calculates sin from latitude of locations 1 and 2 which has been converted to rad sin calculation results location 1 = -0.1107612039 sin calculation results location 2 = -0.1106410154 calculates cos from latitude of location 1, location 2 and longitude difference converted to rad cos calculation results location 1 = 0.9938470484 cos calculation results location 2 = 0.9938604357 cos calculation result longitude difference = 0.999999809 calculates the distance between location 1 to location 2 distance = (sin location 1 * sin location 2) + (cos location 1 * cos location 2 * cos longitude difference) distance =(-0.1107612039 * -0.1106410154) + (0.9938470484 * 0.9938604357 * 0.999999809) distance = 0.9999998039 calculating distance result to acos distance = acos (0.9999998039) distance = 0.00062618 convert acos distance results to degrees distance = 0.03587747121559 converting distance results from degrees to kilometers (km) distance = 0.03587747121559 * 60 * 1.1515 distance = 2,478774486 determine the final result distance = 2.478774486 * 1.609344 distance = 3.989200846 km 4 km the result of the above calculation is 3.989200846 km which is calculated from the starting point of karawang international industrial city and the destination point of griya kost. iii. results and discussion 1. running system analysis the current system analysis aims to see the process of finding a dormitoryin karawang city which is still being carried out by visiting from one dormitoryto another to see the state of the dormitorydirectly. the following is a flowmap on a system that is ongoing or currently running [9]. figure 3 flowmap of the running system 2. proposal system this design aims to meet all the needs of the users of the system to provide a clear and understandable picture. the proposed system describes the system to be built. the system design is made in the form of a flowmap, starting from the user opening the application, until the user manages to get a dormitory. the system design to be built is described in the form of a flowmap as shown in the following figure: 58 | vol.2 no.2, july 2021 figure 4 flowmap of the proposed system 3. usecase diagram usecase diagram is a description of the actors and the needs of the usecase functions needed in the system [10]. the following is an image of the proposed use case diagram: figure 5 usecase diagram of the proposed system the proposed system uses a use case diagram explaining the system to be built with actors such as dormitoryseekers, dormitoryowners, and admins with their respective access. 4. sequence diagram the sequence is generally used to describe a scene or a series of steps taken in response to an event to produce a certain output, as well as what has changed internally and what kind of output has been produced. the following is a modeling of the mobile kost karawang application using a sequence diagram. below is an image of the sequence diagram of the proposed search: figure 6 kost search sequence diagram the picture above is a sequence diagram for the dormitorysearch, in which the user enters the search keyword on the home page, then the system processes the search and the system results will display the search dormitory results taken from the dormitory data. 4. class diagram class diagrams or class diagrams describe the structure of the system in terms of defining the classes that will be created to build the system. with class diagrams, you can make detailed diagrams by paying attention to the specific codes needed by the program and this can be applied to the structure described. the following is a class diagram on a dormitory mobile application. figure 7 class diagram of the mobile dormitory application 5. implementasi location-based service (lbs) location-based service (lbs) is used in the system, especially in conducting searches that can be connected to google maps. the result of location-based service (lbs) is the point from the place or location of the dormitorythat will be stored in the system [11]. 59 | vol.2 no.2, july 2021 lbs implementation begins by finding the last location with latitude and longitude points from the global positioning service (gps) connected to the user's mobile device. this lbs is connected to google maps so that when it is determined the google maps display will match the user's location. 6. implementasi haversine the distance determination is generated from calculations using the haversine formula, the values used are the longitude and latitude points from the starting point to the destination point. the haversine formula begins by calculating the difference between the initial longitude point and the destination longitude point. then the value of the latitude of the starting point and the latitude of the destination and the difference between the longitudes are converted into radians. then calculate sin and cos from the results that have been converted from the value of the starting point and the destination point, but to calculate cos including the results of calculating the difference in longitude of each point [8]. the distance generated from the product of the starting point sin and the destination point sin is added up by the product of the starting point cos and destination point cos and the longitude difference cos. the result is converted to kilometers after finding across and multiplied by 60 * 1.1515 and multiplied by 1.609344 to produce the distance in kilometers [12]. 7. haversine formula calculation analysis the analysis of the haversine formula calculation uses three sample data. the results of calculations with the haversine formula can be seen in the following table: table 1 calculation results of haversine after calculating with the haversine formula, the researcher made an experiment using google measurement to measure the distance. this experiment was made by entering the coordinates of the start point and endpoint in google maps to measure, then using the distance measurement function by drawing a straight line from the starting point to the endpoint that was already available on google maps. this distance calculation experiment can be seen in the following figure: figure 8 google measurement based on the comparison of distance measurements using haversine formula and google maps, there are differences in distance measurement results of 0-10 meters. comparison of distance measurement results as in the following table: table 2 comparison of distance measurement 8. interface design the design of the interface or the interface on the software is used to connect the interaction between the user and the software to be made [13]. interface design or system display is made using pencil software. the interface design of the system building can be seen as shown below [14]. no titik mulai titik tujuan haversine latitude 1 longitude 1 latitude 2 longitude 2 1. -6,359197487 107,2742427 -6,352268604 107,3096558 3.98 km 2. -6,327623336 107,2877780 -6,323338243 107,3012788 1.56 km 3. -6,301365223 107,2780176 -6,304240781 107,2978055 2,19 km no titik mulai titik tujuan haversine google maps latitude 1 longitude 1 latitude 2 longitude 2 1. -6,359197487 107,2742427 -6,352268604 107,3096558 3.98 km 3.99 km 2. -6,327623336 107,2877780 -6,323338243 107,3012788 1.56 km 1.57 km 3. -6,301365223 107,2780176 -6,304240781 107,2978055 2,19 km 2,20 km 1. homepage 2. dormitory details page 60 | vol.2 no.2, july 2021 9. system view implementation after performing the previous stage of analyzing the current system and designing the proposed system, then the process of implementing the system display here is the display of the dormitory mobile application: figure 1 is the home page, on this page, there is a dormitory search column to find a dormitoryby inputting the company name or dormitoryname, then when pressing the search button a list of dormitorys will appear based on the closest distance from the destination company. figure 2 is a page for displaying dormitory in detail such as addresses, categories, ratings or ratings, as well as available dormitory facilities figure 3 is a dormitory booking page, this page displays dormitory orders and can be seen in detail by selecting an order, a page like a figure 4 will appear which contains ordering data and can be inputted by the customer as needed. 10. black box test black box testing is centered on the functional requirements of the software. black box testing makes it possible in software engineering to obtain a set of input, process, and output conditions that are completely in sync with the functionality of a program. the system testing phase is carried out after the system implementation phase is complete. the execution of the test phase is to re-examine all the phases that have been run to find or find errors. the purpose of system testing is to ensure that the system that has been built will function as expected. the following is a table of the results of testing the system interface function using black box testing [15]. table 3 system interface testing no. test name user expected results test result status uji 1. login page dormitor y finder, dormitor y owner, admin the system can display the login page interface when opening the mobile system. the login process can be done by inputting text or logging in via a google account automatically the login page can be displayed. √ 2. homep age cost finder the system can display the home page interface after successful login. the home page can be displayed. √ 3. dormit ory details page cost finder the system can display the dormitory detail page dormitory detail page can be displayed √ 3. order page 4. order details page 1. homepage 2. dormitory details page 3. order page 4. order details page 61 | vol.2 no.2, july 2021 4. search page cost finder the system can display a list of dormitorys that are sought in order based on the closest to the furthest distance the search page can be displayed √ 5. order page cost finder the system can display the order page order page can be displayed √ 6. dormit ory manage ment page dormitor y owner, admin the system can display a dormitory management page, on this page the dormitoryowner can add, view, change and delete dormitorys. admin can view and disable dormitory dormitory manageme nt page can be displayed √ iv. conclusions and suggestions based on the research that has been done and the test results of the mobile dormitory application using locationbased service (lbs) with the haversine formula, it can be concluded that: 1. a dormitory mobile application has been built that can display the location of the closest dormitoryfrom a company in karawang, this method uses the haversine formula to determine the closest distance where haversine is a formula that measures the distance between two points by drawing a straight line between the two points. this formula ignores terrain or obstacles when measuring these two points. 2. the application can provide detailed dormitory information needed by dormitoryseekers such as dormitory addresses, dormitory pictures, available facilities, prices, and other information. reference [1] “karawang kabupaten.” [online]. available: sumber:https://karawangkab.bps.go.id. [2] bayu priyatna and fitria nurapriani, “implementasi koordinat google dan citra kamera pada aplikasi monitoring petugas berbasis android,” buana ilmu, vol. 5, no. 1, pp. 106–121, 2020. [3] g. f. laxmi, f. satrya, and f. kusumah, “perancangan location based service ( lbs ) pada pencarian event aplikasi sahabat jasa berbasis android,” pp. 297–306, 2018. [4] i. a. murdiono, “naskah publikasi perancangan aplikasi mobile location based service (lbs) untuk pencarian lokasi rumah kos di kota sleman berbasis android,” 2020. [5] d. rancadaka, “aplikasi penjualan padi berbasis web dengan menggunakan metode,” vol. 1, no. 1, pp. 29–33, 2020. [6] b. priyatna, s. shofia hilabi, n. heryana, and a. solehudin, “aplikasi pengenalan tarian dan lagu tradisional indonesia berbasis multimedia,” systematics, vol. 1, no. 2, p. 89, 2019. [7] b. huda and b. priyatna, “penggunaan aplikasi content management system (cms) untuk pengembangan bisnis berbasis e-commerce,” systematics, vol. 1, no. 2, p. 81, 2019. [8] h. kusniyati and h. fadhillah, “aplikasi pencarian ustadz untuk wilayah dki jakarta menggunakan algoritma haversine formula berbasis android,” petir, vol. 9, no. 2, pp. 102–111, 2019. [9] h. gunawan and a. k. h. saputro, “pemanfataan aplikasi mobile untuk mempercepat pencarian tempat indekos berbasis android,” j. muara sains, teknol. kedokt. dan ilmu kesehat., vol. 1, no. 2, pp. 85–96, 2018. [10]s. anton, i. c. alex, and n. kristiawan, “sistem penilaian kinerja pegawai dalam pelayanan nasabah pada … (sujarwo dkk.),” penilaian, sist. pegawai, kinerja nasabah, pelayanan kinerja, abstr. lang. unified model., pp. 270–275, 2019. [11]m. ahmad, s. al-ikhsan, and f. satri, “pengembangan aplikasi pencarian wisata kuliner kota bogor berbasis range harga dan metode lbs (location base servise) pada android,” inovatif, vol. 1, no. 2, pp. 1–5, 2018. [12]a. m. abdillah, r. rianto, and n. i. kurniati, “penerapan metode haversine pada aplikasi layanan perbaikan kendaraan berbasis location based service,” juita j. inform., vol. 7, no. 2, p. 81, 2019. [13]b. priyatna and a. hananto, “implementation of application programming interface (api) in indonesian dance and song applications,” systematics, vol. 2, no. 2, pp. 47–57, 2020. [14]gelinas, ulric, oram, alan, wiggins, and william, “accounting information system,” pp. 17–30, 1990. [15]a. ukur, “sistem informasi monitoring kualitas alat ukur,” no. ciastech, pp. 419–428, 2020. paper title (use style: paper title) p-issn : 2715-2448 | e-issn : 2715-7199 vol.1 no.2, 10 july 2020 buana information tchnology and computer sciences (bit and cs) 23 | vol.1 no.2, 10 july 2020 buana information tchnology and computer sciences (bit and cs) the role of financial technology in the creative industries in indonesia nono heryana 1 program studi sistem informasi fakultas ilmu komputer universitas singaperbangsa karawang email: nono@unsika.ac.id rini mayasari 2 program studi sistem informasi fakultas ilmu komputer universitas singaperbangsa karawang email: rini.mayasari@staff.unsika.ac.id ‹β› rudi aprianto 3 program studi sistem informasi stmik pringsewu email: rudiaprianto@gmail.com abstract— in the modern era, the use of technology is growing rapidly to obtain information and various other electronic services, many people who already use technology, especially the internet to facilitate their work. seeing the rapid development of the internet has given rise to various innovations, especially financial technology to meet the various needs of the community including access to financial services and transaction processing. fintech (financial technology) is an innovation in the financial sector that refers to modern technology used to transact, check deposit rates, transfer funds, and perform various other financial services. the creative industries are an industry that utilizes the creativity and skills of individuals use as goods of value. the purpose of this study is to determine the role of financial technology and its constraints in the creative industry in indonesia. the existence of innovation and creativity that arises in the community, making the creative industry is an essential role in the development of a regional economy. the research method used in writing this article is descriptive qualitative. thus, qualitative research only describes responses to situations or events so that it does not explain causality or do hypothesis testing. keywords—financial technology, creative industry, internet, indonesia abstrak—di era modern, penggunaan teknologi berkembang pesat untuk memperoleh informasi dan berbagai layanan elektronik lainnya, banyak orang yang sudah menggunakan teknologi, terutama internet untuk memudahkan pekerjaan mereka. melihat pesatnya perkembangan internet telah memunculkan berbagai inovasi, terutama teknologi keuangan untuk memenuhi berbagai kebutuhan masyarakat termasuk akses ke layanan keuangan dan pemrosesan transaksi. fintech (teknologi finansial) adalah inovasi di sektor keuangan yang mengacu pada teknologi modern yang digunakan untuk bertransaksi, memeriksa suku bunga simpanan, mentransfer dana, dan melakukan berbagai layanan keuangan lainnya. industri kreatif adalah industri yang memanfaatkan kreativitas dan keterampilan yang digunakan individu sebagai barang yang bernilai. tujuan dari penelitian ini adalah untuk mengetahui peran teknologi keuangan dan hambatannya dalam industri kreatif di indonesia. adanya inovasi dan kreativitas yang muncul di masyarakat, menjadikan industri kreatif merupakan peran penting dalam pengembangan ekonomi regional. metode penelitian yang digunakan dalam menulis artikel ini adalah deskriptif kualitatif. dengan demikian, penelitian kualitatif hanya menggambarkan respons terhadap situasi atau peristiwa sehingga tidak menjelaskan hubungan sebab akibat atau melakukan pengujian hipotesis. kata kunci— teknologi keuangan, industri kreatif, internet, indonesia. i. introduction the internet is one of the technologies widely used by humans in modern times like today. ranging from children, teens to adults use it for various purposes such as browsing, chatting, and much more. from this internet emerged various applications and websites that help people in their work [13]. based on the results of the apji and polling indonesia survey in 2018, the number of internet users is 171.18 million. it will be a lucrative market opportunity for applications, systems and technology to reap the domestic market. technology is increasingly developing over time. to make it easier for people to make transactions, check balances and so on, a technology called fintech (financial technology). fintech is a financial service that utilizes technology and software as a forum for the distribution and delivery of information [1]. the background to the emergence of fintech is when there is a problem in society that cannot be served by the financial industry with various obstacles. among them are regulations that are too strict as in the bank and the limitations of the banking industry in serving the community in certain areas [2]. financial technology in bank indonesia regulation number 19/12 / pbi / 2017 is the use of financial system technology that produces new products, services, technology or business models and can have an impact on monetary stability, financial system stability, efficiency, smoothness, security and payment system reliability. financial technology providers which include payment systems, market support, investment management and risk management, loans, financing and capital providers, and other financial services. financial technology or fintech in indonesia is a very potential market opportunity [3]. the fintech companies are mostly micro, small or medium-sized companies that do not have much equity but have a clear idea of how to introduce new services or how to improve existing services in financial service markets [4]. fintech develops in various sectors, ranging from payment startups, lending, financial planning (personal finance), retail investment, financing (crowdfunding), remittances, financial research, and others [2]. like banks, the fintech company's business model also focuses on payment and 24 | vol.1 no.2, 10 july 2020 buana information tchnology and computer sciences (bit and cs) loan services [5]. from the data recorded by the fsa until january 2018 there are 260 thousand people borrowing in the fintech company in indonesia. the presence of financial technology is beneficial for people in accessing financial products and makes it easier to conduct financial transactions with a touch of technology in hand. wherever and whenever people can make transactions without having to come directly to financial companies or waiting in line with various procedures such as banking in general. this can increase financial literacy in indonesia [6]. fintech refers to the integration of finance and technology and is different from electronic finance, where electronic finance continues to technically support the existing financial system [7]. several studies have been conducted previously that discuss the analysis and role of financial technology. one of them is a research conducted by muzdalifa, rahma, & novalia (2018), which discusses the role of fintech in increasing financial inclusion at msmes in indonesia. the results of the study concluded that the presence of a number of fintech companies also contributed to the development of msmes. not only limited to helping finance venture capital, fintech's role has also expanded to various aspects such as digital payment services and financial arrangements [1]. other research has been conducted by chrismastianto (2017) swot analysis discussion related to the installation of financial technology on the quality of banking services in indonesia. research has analyzed the strengths, weaknesses, opportunities and threats (swot) of financial technology installation so that financial technology produces an accurate level of effectiveness in improving the quality of banking services in indonesia so that banking management can apply it to reach all elements of indonesian society [8]. in this study, the author will raise the theme of the role of financial technology in the field of creative industries in indonesia as the subject of the discussion. creative industries are creativity, expertise and talents that have the potential to improve welfare through offering intellectual creation [9]. the government through the trade department has identified the scope of the creative industry covering 14 sectors including advertising, architecture, the art market, crafts, design, fashion, video, interactive games, music, performing arts, publishing and printing, software, television and radio, research and development [10]. the entry of fintech is a new breakthrough in aspects of business in indonesia to become more efficient and comfortable [11], this is what can support the development of fintech in the field of creative industries. creative industries make a significant economic contribution. also, the creative industry creates a favorable business climate and builds the nation's image and identity. on the other hand, the creative industry is based on renewable resources, creates innovation and creativity, which is a competitive advantage of a nation and has a positive social impact [9]. because of the increasingly visible role of financial technology in various fields in modern times like now, the authors will conduct research aimed at analyzing the role of fintech and its constraints in the creative industries in indonesia. ii. method the research method used in writing this article is descriptive qualitative [14]. qualitative research methods are research methods based on postpositivism philosophy, or interpretative and constructive paradigms, which view social reality as something holistic or intact, complex, dynamic, full of meaning and symptom relations are interactive and are used to examine natural conditions of objects, not an experiment, where the researcher as a critical instrument, data collection techniques carried out by triangulation (combined), data analysis is inductive or qualitative and the results of the study emphasize the meaning rather than generalization. this study uses a qualitative method to determine the role and constraints of fintech in the creative industries [15]. qualitative descriptive research is aimed at gathering actual and detailed information, identifying problems, making comparisons or evaluations, and determining what other people are doing in dealing with similar problems and learning from their experiences to establish plans and decisions in the future [4]. thus, qualitative research only describes responses to situations or events, so it does not explain causality or do hypothesis testing. iii. results and discussion in the current era of globalization, the role of financial technology is developing so rapidly for the world economy, one of which is in the field of creative industries. the purpose of this study is to find out the role of fintech and its constraints in the creative industry in indonesia. there are several sub-sectors of the industry or creative economy that have contributed to the country's gdp, as shown in figure 1 below. figure 1. contribution of gross domestic product in creative economy according to subsector (2016)[12] from the picture above it can be seen that the creative economy or industry is a sector that has the potential to be processed more deeply to advance the country's economy. with the entry of fintech is expected to be a breath of fresh air for changes in the creative industry sector. the role of fintech in the creative industry is as follows. 1. capital loan from the data obtained from bekraf (creative economy agency) in 2016 the company's capital results or creative industry businesses are divided into three parts with the following percentage, 92.37% of companies use their capital, 25 | vol.1 no.2, 10 july 2020 buana information tchnology and computer sciences (bit and cs) 24.44% of companies make bank loans, and 0.66% of companies use venture capital. for more details, see in the following figure 2 below. figure 2. access to capital for the creative economy industry 2016[12] from this data fintech can be a solution by providing online capital loans with a simpler and more simple submission process by completing the requirements. several online loan services such as danamas, modalku, and koinworks. 2. digital payment services the fintech company can also provide digital payments that are easier to fund for entrepreneurs or companies, so that entrepreneurs or companies can make payments safely. fintech in this field is ovo, gopay, and jenius. 3. financial management services financial management services are services that provide solutions in the areas of recording expenses, monitoring investment performance and financial consulting at no charge. fintech in this field is a healthy wallet and ngaturduit.com there are several constraints of companies or businesses in the creative economy. according to data obtained from bekraf (creative economy agency), there are several obstacles such as domestic marketing, research and development, physical infrastructure, and others, as shown in the following figure 3 below. figure 3. constraints faced by creative economy enterprises or companies [12] in the field of creative industries in indonesia, the existence of financial technology also has several obstacles, including the following. 1. it infrastructure it infrastructure in indonesia can be said to be uneven, and internet networks can only be felt in big cities like jakarta, bandung, surabaya, and others, while in small cities or remote areas it is still not sure and this is one of fintech's problems in the creative industry. 2. human resources (hr) lack of human resources (hr) in the regions is one of the obstacles to the spread of fintech making it difficult for fintech to develop. human resources (hr) should be educated in the regions to spread fintech evenly. 3. lack of financial literacy many people in rural areas are not familiar with the term fintech about how to use it, its benefits, etc. not knowing this financial literacy makes financial planning and management in the community not suitable. iv. conclusion financial technology (fintech) is an innovation in financial services using technology so that people can easily access financial products and services. fintech can be used to facilitate human being to work, especially in finance such as transactions, checking deposit interest, transferring funds and many other jobs. the presence of several fintech companies contributed to the development of the creative industry in indonesia. in the current era of globalization, the role of fintech is very diverse and expected to help the creative industry players or companies in developing their businesses. there are several roles of fintech that can help businesses or companies in developing their businesses, such as capital loans, digital payment services and financial management services. in helping business people or creative industry companies, fintech also experiences several obstacles such as uneven it infrastructure in indonesia, lack of human resources, and the lack of public financial literacy regarding fintech. references [1] i. muzdalifa, i. a. rahma, and b. g. novalia, “peran fintech dalam meningkatkan keuangan inklusif pada umkm di indonesia (pendekatan keuangan syariah),” j. masharif al-syariah j. ekon. dan perbank. syariah, vol. 3, no. 1, 2018, doi: 10.30651/jms.v3i1.1618. [2] m. ansori, “perkembangan dan dampak financial technology (fintech) terhadap industri keuangan syariah di jawa tengah,” wahana islam. j. stud. keislam., vol. 5, no. 1, pp. 31–45, 2019. [3] a. rusydiana, “bagaimana mengembangkan industri fintech syariah di indonesia? pendekatan interpretive structural model (ism),” al-muzara’ah, vol. 6, no. 2, pp. 117–128, 2018, doi: 10.29244/jam.6.2.117-128. [4] r. muchlis, “schriftenreihe,” at-tawassuth, vol. 3, no. 2, pp. 335–357, 2018, doi: 10.30965/9783846751565_020. [5] r. r. suryono, “financial technology (fintech) dalam perspektif aksiologi,” masy. telemat. dan inf. j. penelit. teknol. inf. dan komun., vol. 10, no. 1, pp. 51–66, 2019, doi: 10.17933/mti.v10i1.138. [6] t. i. f. rahma, “persepsi masyarakat kota medan terhadap penggunaan financial technology (fintech),” at-tawassuth, vol. 3, no. 1, pp. 642–661, 2018. [7] d. kurniawan, e. zusrony, and r. a. kusumajaya, “analisa persepsi pengguna layanan payment 26 | vol.1 no.2, 10 july 2020 buana information tchnology and computer sciences (bit and cs) gateway pada financial technology dengan metode eucs,” j. inf. politek. indonusa surakarta, vol. 4, no. 3, pp. 1–5, 2018. [8] i. adhitya wulanata chrismastianto, “analisis swot implementasi teknologi finansial terhadap kualitas layanan perbankan di indonesia,” j. ekon. dan bisnis, vol. 20, no. 1, pp. 133–144, 2017. [9] a. n. fitriana, i. noor, and a. hayat, “pengembangan industri kreatif di kota batu (studi tentang industri kreatif sektor kerajinan di kota batu),” j. adm. publik, vol. 2, no. 2, pp. 281–286, 2015. [10] l. rahmasari, “pengaruh supply chain management terhadap kinerja perusahaan dan keunggulan bersaing (studi kasus pada industri kreatif di provinsi jawa tengah),” maj. ilm. inform., vol. 2, no. 3, pp. 89–103, 2011. [11] m. g. sitompul, “urgensi legalitas financial technology (fintech): peer to peer (p2p) lending di indonesia,” j. yuridis unaja, vol. 1, no. 2, pp. 68–79, 2018, doi: 10.35141/jyu.v1i2.428. [12] badan ekonomi kreatif indonesia bekraf. “data statistik dan hasil survei khusus ekonomi kreatif.” data statistik dan hasil survei khusus ekonomi kreatif, kerjasama badan ekonomi kreatif dan badan pusat statistik, 8 mar. 2017, www.bekraf.go.id/pustaka/page/data-statistik-danhasil-survei-khusus-ekonomi-kreatif. [13] a. l. hananto and b. priyatna, “rancang bangun aplikasi informasi harga produk pangan dan sembako di pasar kab. karawang,” techno xplore j. ilmu komput. dan teknol. inf., vol. 2, no. 1, 2017. [14] a. m. siregar, s. faisal, y. cahyana, and b. priyatna, “perbandingan algoritme klasifikasi untuk prediksi cuaca,” j. account. inf. syst., vol. 3, no. 1, pp. 15–24, 2020. [15] m. hennink, i. hutter, and a. bailey, qualitative research methods. sage publications limited, 2020. paper title (use style: paper title) p-issn : 2715-2448 | e-issn : 2715-7199 vol.1 no.2, 10 july 2020 buana information tchnology and computer sciences (bit and cs) 27 | vol.1 no.2, 10 july 2020 buana information tchnology and computer sciences (bit and cs) implementation of k-nearest neighbor algorithm for customer satisfaction sutan faisal 1 study program technical information faculty of engineering computer science, university buana perjuangan karawang sutan.faisal@ubpkarawang.ac.id ‹β› nurhayati 2 study program technical information faculty of engineering computer science, university buana perjuangan karawang nurhayati@ubpkarawang.ac.id abstract—customer satisfaction is the company's goal in providing services to its customers. sewa camera cikarang is committed to customer satisfaction. by using the k nearest neighbor (knn) algorithm of this study to analyze customer satisfaction of camera tenants. in this study price, facilities, services and loyalty are input attributes of customer satisfaction. satisfied and dissatisfied is the result of the output. increasing customer satisfaction and increasing profits on cikarang camera rentals is the aim of this research. this study using the knn algorithm obtained accuracy = 98%, recall classification = 86.67%, classification accuracy = 100% and auc = 0.750. it is expected that the results of this study can be used as a reference for building applications that can facilitate companies in obtaining information about customer satisfaction. keywords—datamining, classification, knn algorithm, customer satisfaction. abstrak—customer kepuasan pelanggan merupakan tujuan perusahaan dalam memberikan layanan kepada pelanggannya. sewa kamera cikarang berkomitmen untuk kepuasan pelanggan. dengan menggunakan algoritma k-nearest neighbor (knn) penelitian ini untuk menganalisa kepuasan pelanggan penyewa kamera. dalam penelitian ini harga, fasilitas, layanan dan loyalitas merupakan atribut masukan kepuasan pelanngan . puas dan tidak puas merupakan hasil outputnya. meningkatkan kepuasan pelanggan dan meningkatkan laba pada sewa kamera cikarang adalah tujuan penelitiian ini. penelitian ini dengan menggunakan algoritma knn mendapatkan akurasi = 98%, klasifikasi recall = 86,67%, ketepatan klasifikasi = 100% dan auc = 0,750. diharapkan hasil penelitian ini dapat dijadikan acuan untuk membangun aplikasi yang dapat memudahkan perusahaan dalam memperoleh informasi tentang kepuasan pelanggan. kata kunci— pengumpulan data, klasifikasi, algoritma knn, kepuasan pelanggan. i. introduction a. introduction along with the high level of human activity to meet the needs and needs of daily life, humans need to release their fatigue with a vacation. then it needs to be supported with a camera to capture the moment of his vacation. but not everyone has a camera that is good enough to capture the holidays. public awareness of the elements of service that can be provided by companies is increasing due to advances in education and a more prosperous economy, as well as the development of science and technology. the importance of service quality provided by service companies and in the form of goods is increasingly being realized by consumers. each consumer's assessment of the quality of services / services varies depending on how consumers expect the quality of the service / service based on experience [1]. achieving success in a service business, customer satisfaction must be the basis of management decisions, so management must make increasing customer satisfaction a fundamental goal. in order to provide quality services, the company must continually improve the quality of its human resources and the equipment it leases. this step is important to improve services from time to time. people who judge whether or not the quality of service is called a consumer. by comparing the services they receive with the services they expect consumers can judge the service. consumers who are satisfied with the services provided by a company will make these consumers come back again to use the company's services again. companies that have loyal customers because the company can satisfy their customers. word of mouth promotion without coercion regarding the services it has received will be carried out by loyal consumers [4]. tight competition must be faced by companies in the increasingly rapid development of the business world. the customers he has by the company are expected to be maintained forever. to realize this, it is not something that is easily climatic, as business competition is very tight at the moment considering that there are rapid changes that can occur at any time such as changes in customers, competitors or changes in broad conditions that are always dynamic. this requires policy makers to develop a strategy that is able to achieve sales growth targets, increase the company's market share, and achieve capabilities as the basis for sustainable growth. [1]. the tight competition must be faced by the company in the rapid development of the business world. in general, there are many ways to maintain customers forever, in a very tight mailto:sutan.faisal@ubpkarawang.ac.id mailto:sutan.faisal@ubpkarawang 28 | vol.1 no.2, 10 july 2020 buana information tchnology and computer sciences (bit and cs) business competition it is very difficult to realize it given the many changes that can occur at any time. such as changes in customers. competitors and changes in broad conditions that always change dynamically. this makes policy makers to continue to develop a strategy that can achieve the goals of rental growth, increase market share, and the achievement of capabilities as a basis for sustainable growth [16]. b. definition of data minning data mining is data mining that has long been taken from several series of activities when viewed from the point of view, according to [5]. data mining is an integrated data analysis process that consists of a series of actions based on the definition of the objectives to be analyzed, with data analysis and interpretation of the results. in recent years data mining has attracted the attention of the public and the world of information systems, because useful information in the form of knowledge generated from large data is needed. applications ranging from market analysis, fraud detection, and customer retention, to production control and exploration science are generated from information and knowledge. [7]. according to [5], data mining has the following stages of the process: 1. defining goals for analysis the clearest statement of the problem and the achieved goals are the most important in the correct formulation of the analysis. determining the method to be used is one of the most difficult parts of the process. there must be no room for doubt or uncertainty and clear goals must. 2. selection, organization, and preliminary treatment data the collection or selection of data needed to be analyzed is done after the objectives are analyzed and identified. the ideal source of data is theata's backup company, a "storage room" of historical data that is no longer used. if there is no data storage, the data market can be created by matching different corporate data sources. 3. exploration of data analysis and transforming it at this stage involves an initial exploration analysis of data, which is very similar to thetechnique online analytical process (olap). transformation of the original variables to better understand the phenomena or statistical methods used are carried out at this stage. to highlight anomalous data, different data from other data is used in the analysis of exploration. 4. specifications of statistical methods statistical methods can be used, as well as many available algorithms, so it is possible to classify already available methods. the choice of method used to prepare the analysis depends on the problem being studied or the type of data available. different methods are edited into two main classes according to different stages of data analysis, in particular: a. descriptive method to describe groups of data in a concise manner is the main goal of the method. there is no descriptive hypothesis between the available variables. included in this group are the association method, loglinearmodel, graphical model). b. prediction method the purpose of this class method is to describe one or more variables that are performed by finding classification or prediction rules based on the data. these rules help to predict or classify one or more answers or future variables of the target variables in relation to what is happening with the explanatory or input variable. included in this method are neural networks, decision trees, and linear and logistic regression models. 5. data analysis based on the method chosen, which will then be applied to the statistical method to be used then translating into the appropriate algorithm to get the required results based on available data. 6. evaluation of the methods used and comparison for the analysis of the final model selection 7. commentary on the selected model and its use in the decision-making process. c. clasification and prediction classification and prediction is a method that can make smart decisions. researchers have now proposed a number of classifications and forecasting methods for machine learning, pattern recognition, statistical research. in this study, we focus on classifying methods in data mining as part of the machine learning process. the form of data analysis that can be used to extract models to predict future trends in data to be predicted is the classification and prediction of data mining. the classification process is divided into two stages, first the learning process in which the classification algorithm is used to analyze training data. is, the results of the presentation of the learning model or classifier in the form of classification rules, the two phases of the classification process, estimating the accuracy of the classification model or classifier from the test data. if the accuracy is accepted, the model is applied to find out the predicted results of new data. bayesian methods, bayesian networks, algorithm-based rules, neural networks, vector machine support, mining rules associations, k-nearest neighbors, case-based reasoning, genetic algorithms, rough sets and fuzzy logic are the classification techniques used. focusing the nearest neighbor (knn) k algorithm in this study. d. data minning methods the idea of people already having knowledge in the process of classifying management has already been widely used. but talking about taxonomy (tassein = classify + nomos = science, law) its use as a science of grouping living organisms (alpha taxonomy) at first ,has since become a general science group, including the principle of classification (taxonomic schemes). thus, classification (taxonomy) processes the placement of an object (concept) based on a number of categories, each object (concept) based on ownership. [6]. four basic components for the classification process: 1. class: the dependent variable of the model is the categorical variable to represent the 'label' that uses the object after its classification. examples of lessons are: heart attack, customer loyalty, stellar lesson (galaxy), earthquake lesson (storm), etc. 29 | vol.1 no.2, 10 july 2020 buana information tchnology and computer sciences (bit and cs) 2. predictors: classification of data and based on the classification made from the model represented by the characteristics (attributes) which are independent variables. examples of such predictors are: smoking, drug consumption, blood sugar, sales frequency, sex status, satellite images, geological record, and wind speed direction, season, etc. 3. training dataset the data used for the 'training' model to recognize according to class, based on predictions available from the two two component data values before. 4. testing the dataset: contains new data classifications based on the model built on, and classifications that are accurate (model performance) so they can be evaluated [6]. a. there are no other attributions in the separate post b. there are no records inbranch an empty e, k neighrest neighbor k-nearest neighbor (knn) is included in the instancebased learning group. this algorithm is also one of thetechniques lazy learning. knn searches the k group of objects in the training data that is closest (similar) to the object in new data or testing data (similar) to the object in new data ordata testing [15]. case in point, for example it is desirable to find a solution to the problem of a new patient by using a solution from an old patient. to find solutions from new patients, closeness to old patient cases is used, solutions from old cases that have closeness to new cases are used as a solution. there were new patients and 4 old patients, namely p, q, r, and s (figure 2). . when there is a new patient, the solution is taken from the case of the elderly patient who has the greatest kinship. fig. 1 ilustrasi knn for example, d1 distance between new patients and patient p, d2 distance between new patients and sick q, d3 distance between new patients and sick r, d4 distance between new patients and patient s. the picture shows that d2 is closest to the new case. thus, the patient q solution will be used as a solution for the new patient. (henny leidiyana, 2013) euclidean distance and manhattan distance (city block distance) are ways to measure the proximity between new data and old data (training data), the most commonly used is euclidean distance. [2], namely: where a = a1, a2, ..., an, and b = b1, b2, ..., bn represents the n attribute values of the two records. for attributes with category values, measurements with euclidean distance do not match. instead, the following functions are used [10]: different (a, b) {0 if ai = bi = 1 besides where ai and bi are the category values. if the attribute value between the two records being compared is the same, the distance value is 0, the meaning is similar, on the contrary, if it is different then the value of proximity is 1, it means it is not similar at all. for example the color attribute with red and red values, the value of proximity is 0, if red and blue then the value of proximity 1. normalization is done if measuring the distance from attributes that have large values, such as income attributes. normalization can be done with min-max normalization or z-score standardization [10]. if thedata training consists of a mixture of numerical and category attributes, the use of min-max normalization is preferred [10]. to calculate the similarity of cases, a formula is used [9]: note: p = new cases q = cases in storage n = number of attributes in each case i = individual attributes between 1 to n f = function similarity attribute i between cases p and case q w = weight given to i attribute e. evaluation and validation of data mining prediction methods in this study cross validation, confusion matrix, and roc (curvesreceiver operating haracteristic) curve methods are used for evaluation and validation. 1. cross validation to predict the error rate standard testing is done. in getting the overall error rate, the training data is randomly divided into several parts with the same comparison then the error rate is calculated section by section, then calculate the average for all error rates 2. confusion matrix table 2.1 is the method used, one class is considered positive and the other negative, if the dataset consists of only two classes. percentage of accuracy of data records that are classified correctly after testing the classification results is the result of evaluation with a confusion matrix that has accuracy, precison, and recall.accuracy values [7]. the proportion of positive predicted cases that are also true positive on the actual data is called precision or confidence. the proportion of true positive cases that is correctly predicted correctly is called recall or sensitivity. [12]. table 1 model conflusion matrix correct classification classified as + + true positives false negatives false positives true negatives 30 | vol.1 no.2, 10 july 2020 buana information tchnology and computer sciences (bit and cs) true positive is the number of positive records that are classified positively, false positive is the number of negative records that are classified positively, false negative is the number of positive records classified as negative, true negative is the number of negative records classified as negative, then enter the test data. to get the amount of sensitivity (recall), specifity, precision, and accuracy enter the value of the test data into the confusion matrix. sensitivity is used to compare the number of t_pos to the number of positive records, while the comparison of the number of t_neg to the number of negative records is used precision. the equation below is used to calculate it 7]: sensitifity = 𝑡_𝑝𝑜𝑠 𝑝𝑜𝑠 (3.0) specifity = 𝑡_𝑛𝑒𝑔 𝑛𝑒𝑔 precision = 𝑡_𝑝𝑜𝑠 𝑡_𝑝𝑜𝑠+𝑓_𝑝𝑜𝑠 accuracy = sensitivity pos + (pos+neg) specifity neg (pos+neg) remarks: t_pos = number of true positives t_neg = number of true negative p = number of record positives n = number of tuples negatives f_pos = number of false positives 3. roc curve accuracy and visually comparing classifications can be demonstrated by the roc curve. confusion matrix specified by the roc. two-dimensional graphics with horizontal lines as false positives and vertical lines as true positive are called roc (vercellis, 2009). to measure the difference in performance the method used is generated from the calculation of the area under curve (auc). the formula used by auc θr = 1 mn ∑ ∑ ψm i=1 n j=1 (xtr, xjr) where : 𝟁(x,y) = { 1 𝑌 < 𝑋 1 2 𝑌 = 𝑋 0 𝑌 > 𝑋 description: x = positive output y = negative output ii. method in this study using rapidminer studio 9.0 testing tools, using the following methodology: fig. 2 methodology used a. dataset is a collection of data, a database table represented by a dataset, or it could be a data matrix where each particular variable is represented by a column, the amount of data is represented by a row. the retrieve operator loads the rapidminer object into the process used in this research. exampleset, but can also be a collection or a model. data is retrieved this way as well as meta data from the rapidminer object. b. validation the operator used to perform simple validation randomly divides exampleset into a training set to set the test and evaluate the model. split validation to estimate the performance of the learning operator (usually in an invisible data set) is performed by this operator. in practice, it will be shown how accurate a model estimate (learned by certain learning operators). c. knn algorithm in this study an experiment was carried out using the classification method of decamination tree datamining knn algorithm on customer satisfaction questionnaire data on cikarang camera rental. data will be processed using the knn algorithm and produce a model, then the resulting model will be tested cross validation which produces accuracy, precision, recall and auc. d. apply model learning algorithm which is the first model trained on exampleset by other operators. after that, this model can be applied to another exampleset called apply model. to get predictions on data that are not visible or to transform data by applying the preprocessing model is the goal of applying the model. the model attribute must be compatible with exampleset where the model is applied. exampleset apply the model must have the same number, sequence, type, and role attributes as exampleset used to generate the model. e. performance the the operator is used to evaluate the statistics of a binomial classification task, ie a classification task whose 31 | vol.1 no.2, 10 july 2020 buana information tchnology and computer sciences (bit and cs) label attribute has a binomial type. this operator provides a list of performance criterion values from the binomial classification task. measure the results of this study using a confusion matrix (accuracy, recall classification, classification accuracy) and the roc curve. iii. results and discussion a. data set analysis results the data set used is the customer satisfaction questionnaire data set for cikarang camera rental, this data set contains data information about customer satisfaction questionnaires regarding prices, facilities, services and loyalty. the total data in this data set is 100 records, each of which has 10 attributes including: 1. no (integers, roles: id) 2. name (polynominals) 3. price x1 (integers) 4. facility x2 (integers) 5. services x3 (integers) 6. loyalty x4 (integers) 7. results (binomials: satisfied & dissatisfied) from the attribute data set above (1 to 7) the training & test process will be carried out using the 10 fold cross validation method, while the 7th attribute will be the target of the results of the classification process. and here we will try to analyze the difference between accuracy and error obtained by comparing the predicted results and results. fig. 3 10 fold cross validation the questionnaire can be illustrated below: fig. 4. questionnaire form data that has been processed using ms excel fig. 5. data questionnaire that has been processed b. experiment and evaluation results in this experiment there are 7 attributes which will be trained and 2 values that indicate the target (classification) on the 7th attribute, which means the knn algorithm is initialized, 6 input attributes and 1 output attribute. the results of this study are: 1. confusion matrix table number of true puas (tp) is 124 records classified as true positive 124 records and false negative (fn) of 0 records . next 26 records for true dissatisfaction (ttp) are classified as true positive 23 records and false negative as many as 3 records. 2. pervormance vector no nama harga x1 fasilitas x2 pelayanan x3 loyalitas x4 hasil 1 anwar 6 4.33 4.00 3.50 puas 2 maulana 6 4.33 4.00 3.50 puas 3 budiman 5 4.33 3.75 3.50 puas 4 geofany 5 4.33 3.75 3.00 puas 5 fiki ananda 5 4.33 3.75 3.00 puas 6 haryanto 2 3.00 3.00 2.50 tidak puas 7 rizky narezka 5 4.33 4.25 3.75 puas 8 aditia 5 4.00 4.50 3.75 puas 9 agil 5 4.33 3.75 4.00 puas 10 fadilah 3 3.00 3.75 2.75 tidak puas 11 purwati 5 4.33 3.50 3.75 puas 12 nurhajjah 5 3.67 3.50 4.25 puas 13 umay 5 4.67 4.00 3.50 puas 14 jesica 3 3.67 3.25 2.75 tidak puas 15 krismonga 5 4.00 3.50 3.50 puas 16 marzuki 3 3.00 2.75 2.75 tidak puas 17 akbar 3 3.67 2.50 2.50 tidak puas 32 | vol.1 no.2, 10 july 2020 buana information tchnology and computer sciences (bit and cs) 3. roc curve iv. conclusion the classification method using the knn algorithm is very good for determining the correctness of classification in data mining. evidenced by the results of accuracy = 98%, classification recall = 86.67%, classification precision = 100% and auc = 0.750. references [1] abdul rohman, model algoritma k nearest neighbour (knn) untuk prediksi kelulusan mahasiswa, universitas pandanaran semarang, 2015. [2] bramer, max. principles of data mining. vol. 180. london: springer, 2007. [3] basuki, achmad dan syarif, iwan. 2003. modul ajar decision tree. surabaya : pens-its. [4] deddy setyawan, “analisis kepuasan pengguna jasa transportasi taksi untuk meningkatkan loyalitas,” universitas diponegoro, 2010. [5] giudici, paolo, and silvia figini. front matter. john wiley & sons, ltd, 2009.applied data mining for business and industry. [6] gorunescu, florin. data mining: concepts, models and techniques. vol. 12. springer science & business media, 2011. [7] han j, kamber m. 2001. data mining : concepts and techniques. simon fraser university, morgan kaufmann publishers. [8] henny leidiyana, 2013. penerapan algoritma k nearest neighbor untuk penentuan resiko kredit kepemilikan kendaraan bermotor. jurnal penelitian ilmu komputer, system embedded & logic. [9] kusrini&luthfi,e.t. 2009. algoritma data mining. yogyakarta : andi publishing. [10] larose, d.t, 2006. discovering knowledge in data: an introduction to data mining. john willey &sons, inc. [11] m rizki ilham, purwanto. 2016. implementasi datamining menggunakan algoritma c 4.5 untuk prediksi kepuasan pelanggan. udinus semarang. [12] powers, david martin. "evaluation: from precision, recall and f-measure to roc, informedness, markedness and correlation." (2011). [13] rapid-i gmbh. (2008).rapidminer-4.2-tutorial. germany: rapid-i. [14] resty mardiana, “faktor – faktor yang memperngaruhi kepuasan pengguna jasa taksi blue bird,” jakarta, universitas gunadarma, 2010. [15] sachdeva, m., zhu, s., wu, f., wu, h., walia, v., kumar, s., ... & mo, y. y. (2009). p53 represses cmyc through induction of the tumor suppressor mir145. proceedings of the national academy of sciences, 106(9), 3207-3212. [16] tan s, kumar p, steinbach m. 2005. introduction to data mining. addison wesley. p-issn : 2715-2448 | e-isssn : 2715-7199 vol.3 no.1 january 2022 buana information tchnology and computer sciences (bit and cs) 22 | vol.3 no.1, january 2022 sales system using apriori algorithm to analyze consumer purchase patterns elfina novalia1 study program information systems faculty of engineering and computer science, universitas buana perjuangan karawang elfinanovalia@ubpkarawang.ac.id apriade voutama2 study program information systems faculty of computer science, universitas singaperbangsa karawang apriade.voutama@staff.unsika.ac.id syahri susanto3 school of oil technic akademi minyask & gas balongan syahri28@gmail.com ‹β› abstract—penelitian ini bertujuan untuk membuat sebuah sistem penjualan untuk mendapatkan data pesanan tepat waktu, tidak terlambat sampai dalam hitungan hari, dan data menjadi terstruktur. serta mengembangkan solusi untuk mengolah data transaksi penjualan yang akan semakin banyak menggunakan algoritma apriori untuk mengetahui pola pembelian konsumen sehingga dapat menjadi output untuk pengambilan keputusan atau pengetahuan. penelitian ini menggunakan metode kualitatif untuk memperdalam pemahaman tentang fenomena yang ada saat ini sedalam mungkin. hal ini menunjukkan pentingnya kedalaman dan detail dari data yang dipelajari. pengembangan sistem menggunakan metode waterfall karena sangat sesuai dengan kebutuhan sistem yang akan dibangun. dari hasil penelitian, perhitungan sampel data transaksi dengan total 12 data pada tanggal 7-8 agustus 2021 menggunakan alat tanagra menghasilkan aturan asosiasi bahwa jika anda membeli pusaran, anda akan membeli caraco dengan nilai support sebesar 58% dan nilai confidence 100%, memiliki nilai lift ratio sebesar 1,3 menyatakan bahwa kedua produk tersebut memiliki keterikatan yang kokoh satu sama lain. diikuti oleh jika anda membeli faraco, anda akan membeli pusaran. jika anda percaya pada kristal, anda akan membeli arco yang memenuhi kriteria parameter yang ditentukan dengan nilai dukungan minimum 20% dan kepercayaan minimum 50%. keywords: penjualan, data mining, apriori algorithms. abstract—this study aims to create a sales system to get order data on time, not too late to result in days, and the data becomes structured. as well as develop solutions to process sales transaction data which will increasingly use a priori algorithms to find out consumer buying patterns so that they can be output for decision making or knowledge. this study uses a qualitative method to deepen understanding of the phenomena currently happening as profoundly as possible. this shows the importance of depth and detail of the data studied. the system development uses the waterfall method because it fits perfectly with the needs of the system to be built. from the results of the study, calculating a sample of transaction data with a total of 12 data on august 7-8, 2021, using the tanagra tools resulted in a rule association that if you buy a vortex, you will buy a caraco with a support value of 58% and a confidence value of 100%, having a lift ratio value of 1.3 stated that the two products have a solid attachment to each other. followed by if you buy faraco, you will purchase a vortex. if you believe in a crystal, you will buy an arco that meets the specified parameter criteria with a minimum support value of 20% and minimum confidence of 50%. keywords: sales, data mining, apriori algorithms. i. introduction technological developments from time to time continue to develop very quickly. likewise, what happened in the industrial era 4.0, where we are currently in an age that can be made easier to find and get what information we need through cyberspace as if the world is in our own hands. technology brings significant changes to its users. technology has both positive and negative impacts. in the industrial world, technology has an essential role in maintaining and caring for important data owned by a company agency. distributor ud bangun persada is a subsidiary of pt triton paint located in malang city, east java. pt triton paint is a company engaged in paint production. various kinds of paints produced include wall paint, iron paint and wood. distributor cat ud bangun persada, located in the karawang area, is a subsidiary of the west java main office in the cirebon area. currently, it has six employees consisting of 1 sales head, one admin, three sales and one courier. as well as having material shops that are members of the ud bangun persada paint distributor, this number will continue to increase from time to time. areas that become distribution centres for ud distributors build persada such as karawang, purwakarta, bekasi and other areas. ud bangun persada paint distributor only markets paint products produced by pt triton paint. in addition, it does not market products from other companies. in the case of sales transactions, they still record in writing, namely by way of sales recording what is ordered by the consumer on the order form. then the structure is given to the admin or hung on the shelf. after that admin inputs consumer orders, several problems are found, including the sales data being not on time, and sales data is still often carried by sales and not stored in a structured manner. from sales data that is increasingly piling up, a solution can be made to be processed as well as possible using apriori algorithm data mining to find out consumer purchasing patterns, namely the itemset pattern 23 | vol.3 no.1, january 2022 to determine the attachment of one item pattern to another item that can produce information that can be used for decision making. decisions and gain knowledge. ii. method a. system a system is a procedural network of interconnected ones that collectively perform operations or achieve certain goals [1]. the system is a procedure or interrelated elements that have input, process, and output for the system to achieve its goals [2]. b. information information is data that has been classified, processed or interpreted for use in decision making. information processing systems convert data into information or process unnecessary data to be useful to the recipient[3]. information is data that is processed in a format that is more useful and meaningful to the recipient, and data is a source of information that describes actual events [4] c. sales sales are receipts obtained from the delivery of merchandise or from the delivery of services on the stock exchange as consideration items, namely in the form of cash, cash equipment or other assets [5]. sales is the gathering of a buyer and seller with the aim of exchanging goods and services based on valuable considerations, such as money considerations [6]. d. apriori algorithm the apriori algorithm is a method for finding the pattern of relationships between one or more elements in a data set, the apriori algorithm is known as the market basket. the a priori algorithm can understand the buying patterns of consumers in the case that there is a 50% chance that consumers will buy goods a and b and then goods and c [7]. apriori is a class algorithm that helps learn association rules. it works against transactions. the algorithm tries to find a common subset of a data set. a minimum threshold must be met for the association to be confirmed[8]. e. tanagra tanagra is free software for academic and research purposes. this research involves several methods in data mining ranging from data exploration analysis, statistical learning, machine learning to databases [9]. tanagra is one of the data mining software in which several data mining methods are provided, starting from exploring data analysis, statistical learning, machine learning and databases. unlike most data mining software, tanagra is an open source based software where everyone can access the source code, and add their own algorithms, as long as he agrees and conforms to the software distribution license.[10]. f. waterfall the waterfall model is the simplest sdlc (software development life cycle) model. this model is only suitable for software development with specifications that do not change [1]. waterfall provides a sequential or sequential software lifeflow approach starting from the analysis, design, coding, testing and support stages[11]. g. data collection techniques this method is compiled based on the results of the analysis of the research model that will be used, the results of the selection of system development, the waterfall model is used [12]. literature study is done by looking at books, journals and previous scientific works to learn and find out information related to the author's research. observation or observation is one of the primary data collection techniques by directly observing an activity carried out [13]. interviews are primary data collection techniques by directly face to face with the interviewee[13]. h. system development method figure 1 research flowchart analysis of the data to be searched, such as supporting data for system development and observing how the sales process and sales transaction data processing, as well as requiring sales transaction data samples for the a priori algorithm calculation process using tanagra tools which can find out patterns of consumer buying associations, and references these calculations will be compared with the a priori algorithm calculations in the system. the need for transaction data samples with the aim of understanding the attributes on the sales form and selecting attributes for the purpose of data mining processes. in the data analysis stage, the sample of transaction data is calculated using the tanagra tools with a minimum reference of 20% support and 50% minimum confidence. the result is the conclusion of the rule association that is formed from the calculation of transaction data using the tanagra tools. the design stage is the design of the system to be built. this design uses the uml (unified modeling language) model, consisting of several steps, such as use case diagrams, activity diagrams, sequence, and class diagrams. implementation from the design stage into coding in a programming language. the programming language used is php java and uses a mysql database and the codeigniter framework. whitebox testing is focused on the internal system, namely the source code of the program [14]. blackbox testing is done by testing system applications that involve users, which aims to find out the shortcomings of the application system that has been built. maintenance at this stage, the care that has been developed is carried out. this treatment is to prevent errors found in the system carried out on the method according to the software requirements to keep it stable[15]. iii. results and discussion a. data analysis the data analysis process uses the tanagra tools, in the early stages of determining a sample of transaction data. to find the rule association, determine the minimum support with 20% criteria while the minimum confidence is 50%. 24 | vol.3 no.1, january 2022 figure 2 sample transaction data before using the tanagra tools, in the next stage, converting transaction data as shown above into binary format with the following results: figure 3 tabular format after converting the data into binary format, the next step is to upload the binary format data into the tanagra tool. figure 4 tanagra dashboard then in the next process click define status 1, select the attribute that will be processed for mining, then select ok. figure 5 define status 1 then at the bottom select the association menu, then drop frequent itemsets into define status 1. in frequent itemsets 1, right click, then select parameters. as has been determined in the early stages of the parameters for a minimum support of 20%. then select ok. figure 6 frequent itemsets 1 the next step is right click on frequent items 1. click execute, right click on frequent itemsets 1 then select view. then the itemset minimum support 20% will appear as follows. figure 7 results of 20% support itemsets then select the association menu again and select a priori, then drop into define status 1. right click on a priori, select parameters then input minimum support 20% and minimum confidence 50% then press ok. figure 8 input support and confidence value the next step right click on a priori. press execute, right click then view. then the results of the rule association will appear with a minimum support of 20% and a minimum of 50% confidence as follows. figure 9 results of the rule association from the results of mining calculations using tanagra tools with minimum support parameters of 20% and 25 | vol.3 no.1, january 2022 minimum confidence of 50%, conclusions can be drawn that produce association rules. if you buy a vortex, you will buy a varaco with a support value of 58% and a confidence value of 100%. having a lift ratio value of 1.3 indicates that the two products have a strong attachment to each other. followed by varaco => vortex, crystal => arco that meets the predetermined parameter criteria. b. system design the system design is the result of the implementation of the system analysis stage, in system modeling using uml (unified modeling language) diagrams which consist of use case diagrams, activity diagrams, sequence diagrams and class diagrams, the following are the results of the system design analysis. a) use case diagram use case admin the admin use case describes the role of actors who have access to the system, such as having access rights to login, managing customer data, managing product data, managing transactions, managing transaction data for data mining processing purposes, managing mining processes to generate association rules, managing results. mining, reports and can manage user management. the sales use case describes the role of actors who have access to the system, such as having login access rights, and managing sales transactions for product orders from customers. use case sales coordinator describes the role of actor activities who have access to the system, and have login access rights. the sales coordinator can manage transaction data for mining processing purposes, manage the mining process to generate association rules, and manage reports. figure 10 use case diagram b) activity diagram this activity is carried out by the admin to process sales transaction data using the apriroi algorithm, the admin selects the a priori submenu and then the processing mining form appears. after that select the transaction date range to be processed then input the minimum support and minimum confidence then submit the process, then the system displays the association rule that has been processed by mining. c) sequence diagram this activity is for admins when doing the sales transaction data mining process. figure 11 apriori process sequence diagram d) class diagram class diagram is a description of the system structure consisting of database attributes used in building a system, the following is a description of the class diagram. figure 12 class diagram a. system implementation the following is the association rule resulting from the system calculation. figure 13 rule association counting system the following is the association rule resulting from the calculation of the tanagra tools. figure 14 rule association tanagra 26 | vol.3 no.1, january 2022 every user who has logged in according to their access rights will be immediately directed to the dashboard page. figure 15 dashboard pages the transaction page is a process page for finding association rules, this page can be accessed by admin and sales coordinators. figure 16 sales transaction page mining process page is a process page to find association rules, this page can be accessed by admin and sales coordinator. figure 17 mining process pages the mining results page is a history of the results of the association rule search, this page can be accessed by admin and sales coordinators figure 18 mining results page iv. conclusion the results of calculations using the tanagra tool with a total of 12 transaction data from august 7-8, 2021, produce a rule association: if you buy a vortex, you will buy a caraco with a support value of 58% and a confidence value of 100%. a lift ratio value of 1.3 indicates that the two products have a solid attachment to each other. you are followed by caraco => vortex, crystal => arco that meets the predetermined parameter criteria. after conducting an experimental calculation using the tanagra tool and then comparing it with a system calculation and producing almost similar measures, the sales system using the a priori algorithm data mining follows the needs. so that ud bangun persada distributors can find out the product is purchasing patterns of their members in the sales application. then the sales application generates sales data in real-time, and the sales data becomes structured so that you can view reports as needed. reference [1] a. suryanto, “rancang bangun sistem informasi pendaftaran artis berbasis web menggunakan model waterfall (studi kasus : team management agensi),” jurnal khatulistiwa informatika, vol. iv, no. 2, pp. 117–126, 2016. [2] i. laengge, h. f. wowor, and m. d. putro, “sistem pendukung keputusan dalam menentukan dosen pembimbing skripsi,” jurnal teknik informatika, vol. 9, no. 1, 2016. [3] fitri ayu and nia permatasari, “perancangan sistem informasi pengolahan data pkl pada divisi humas pt pegadaian,” jurnal infra tech, vol. 2, no. 2, pp. 12–26, 2018. [4] r. jermias, “analisa sistem informasi akuntansi gaji dan upah pada pt. bank sinarmas tbk. manado,” jurnal riset ekonomi, manajemen, bisnis dan akuntansi, vol. 4, no. 2, pp. 814–828, 2016. [5] p. t. smart, “analisis persediaan dan penjualan terhadap arus kas operasi pada,” vol. 15, no. 1, 2021. [6] e. hikmawati, “penjualan obat pada apotek pusat dan cabang,” no. 8, pp. 1–8. [7] m. syahril, k. erwansyah, and m. yetri, “penerapan data mining untuk menentukan pola penjualan peralatan sekolah pada brand wigglo dengan 27 | vol.3 no.1, january 2022 menggunakan algoritma apriori,” jurnal teknologi sistem informasi dan sistem komputer tgd, vol. 3, no. 1, pp. 118–136, 2020. [8] i. wahyudi, s. bahri, and p. handayani, “aplikasi pembelajaran pengenalan budaya indonesia,” vol. v, no. 1, pp. 135–138, 2019. [9] f. rahmawati and n. merlina, “metode data mining terhadap data penjualan sparepart mesin fotocopy menggunakan algoritma apriori,” piksel : penelitian ilmu komputer sistem embedded and logic, vol. 6, no. 1, pp. 9–20, 2018. [10] h. widayu, s. d. nasution, n. silalahi, and mesran, “data mining untuk memprediksi jenis transaksi nasabah pada koperasi simpan pinjam dengan algoritma c4.5,” media informatika budidarma, vol. vol 1, no, no. 2, p. 37, 2017. [11] s. supriyati and d. m. rizky, “model perancangan sistem informasi akuntansi budidaya perikanan berbasis sak emkm dan android,” is the best accounting information systems and information technology business enterprise this is link for ojs us, vol. 3, no. 2, pp. 301–315, 2018. [12] b. huda and b. priyatna, “penggunaan aplikasi content management system (cms) untuk pengembangan bisnis berbasis e-commerce,” systematics, vol. 1, no. 2, p. 81, 2019. [13] a. l. hananto and b. priyatna, “rancang bangun aplikasi informasi harga produk,” technoxplore jurnal ilmu komputer & teknologi informasi, vol. 2, no. 1, pp. 10–20, 2017. [14] s. kasusdi, p. t. rumah, and s. padjadjaran, “information system journal,” pp. 1–10, 2010. [15] d. rancadaka, “aplikasi penjualan padi berbasis web dengan menggunakan metode,” vol. 1, no. 1, pp. 29– 33, 2020. p-issn : 2715-2448 | e-isssn : 2715-7199 vol.3 no.2 july 2022 buana information tchnology and computer sciences (bit and cs) 53 | vol.3 no.2, july 2022 changing data image into numeric data on kiln manufacture machinery use optical character recognition (ocr) agustia hananto1 study program information system universitas buana perjuangan karawang agustia.hananto@ubpkarawang.ac.id elfina novalia 2 study program information system universitas buana perjuangan karawang elfinanovalia@ubpkarawang.ac.id ‹β› goenawan brotosaputro3 study program master in computer science universitas budi luhur goenawan.brotosaputro@budiluhur.ac.id abstrak— pada proses pembuatan genteng keramik terdapat proses pembakaran menggunakan mesin kiln manufacture atau oven. untuk memastikan proses pembakaran berjalan dengan baik, dilakukan pemantauan 39 parameter yang ada pada mesin kiln tersebut yang harus di awasi secara manual berdasar data yang dihasilkan citra pada mesin kiln tersebut. proses pemantauan parameter tersebut tidak efektif yang disebabkan oleh kesalahan atau kelalaian manusia dan sifat manusia lainnya yang mengakibatkan kerugian bagi perusahaan. oleh karena itu, dibutuhkan sistem untuk menyimpan data parameter mesin pembakaran genteng keramik yang dapat disimpan kedalam sebuah data base. setelah dilakukan analisis pada perangkat pembakaran (kiln) tersebut, recorder sensor yang menampilkan data parameter dapat diakses melalui jaringan lan (local area network), akan tetapi data yang dihasilkan dalam bentuk citra bukan dalam bentuk data digital alfanumerik. data citra yang didapat perlu diterjemahkan menjadi data alfanumerik sebagai sumber data. melalui pengenalan optical character recognition (ocr) dengan metode template matching, citra tersebut diubah menjadi data alfanumerik sehingga dapat di simpan dalam sebuah data base. dari hasil penelitian ini, prototype sistem yang dibuat mendapatkan akurasi sebesar 100.00% untuk konversi data citra ke data alfanumerik, kata kunci: citra digital, data numerik optical character recognition, kiln manufacture. abstract— in the process of making ceramic tiles, there is a combustion process using a kiln manufacture or oven. to ensure the combustion process goes well, 39 parameters are monitored on the kiln engine which must be monitored manually based on the data generated by the image on the kiln engine. the process of monitoring these parameters is ineffective due to human error or negligence and other human traits that result in losses for the company. therefore, a system is needed to store the parameter data of the ceramic tile combustion engine which can be stored in a database. after analyzing the kiln, the sensor recorder that displays parameter data can be accessed via a lan (local area network), but the data generated is in the form of an image, not in the form of alphanumeric digital data. the image data obtained need to be translated into alphanumeric data as a data source. through the introduction of optical character recognition (ocr) with the template matching method, the image is converted into alphanumeric data so that it can be stored in a database. from the results of this study, the prototype system made obtained an accuracy of 100.00% for the conversion of image data to alphanumeric data, keywords: digital image, optical character recognition numerical data, kiln manufacture. i. introduction pt. xyz is a leading manufacturer of glazed ceramic tile (and its accessories). one of the manufacturing processes for these products is the combustion process at a temperature of 1100 degrees celsius using a kiln manufacture machine or oven so that it can produce quality and durable ceramic tiles. in the process of burning ceramic tile products, direct monitoring is carried out by employees for 24 hours by observing 39 data parameters displayed through images on the monitor on the kiln engine panel. in the process of monitoring the kiln manufacture (oven) machine, the sensor data will be displayed in the form of an image that appears on a screen that will be updated every 5 seconds which must be monitored during the process. the image displayed on the oven screen will then be recorded manually and really must be monitored directly. however, human endurance and physical condition greatly affect the results of the monitoring. based on the initial analysis of the problem, it can be concluded that the kiln manufacture (oven) machine only displays 39 parameter data on a screen in the form of an image that will be updated every 5 seconds, the data is stored on a small capacity record machine and cannot be accessed by data. to overcome this problem, it would be possible to create a system or tool that can convert the image into alphanumeric data that can provide real-time information (monitoring automation). technological developments are increasingly developing more advanced than before, as well as image processing technology or digital images. image processing is a method of processing images (images / images) into digital form for certain purposes. one of the digital image processing studies is optical character recognition (ocr) which is a character recognition process through preprocessing, segmentation, feature extraction and recognition, optical character recognition (ocr) is one of the study areas of pattern recognition (pattern recognition) in digital images that classifying or describing an object based on quantitative measurements of its main features or properties. with this method is expected to provide a solution to pt. xyz in overcoming the problems that exist in the kiln manufacture (oven) machine. 54 | vol.3 no.2, july 2022 the image can be accessed via a local area network (lan) and through the introduction of optical character recognition (ocr) the image can be converted into data in alphanumeric form as needed and can be stored into a system that can process the data and it is hoped that the application can provide information automatically. quickly and precisely to the user or user. ii. method 2.1 . study of literature a. image processing image processing is the process of processing pixels in a digital image for a specific purpose. initially, image processing was carried out to improve image quality, but with the development of the computing world, which is marked by the increasing capacity and speed of computer processing and the emergence of computational sciences that allow humans to retrieve information from an image[3]. the image processing process is a diagrammatic process starting from image retrieval, image quality improvement, up to a representative statement of the imaged image can be seen in the figure 1. figure 1 diagram of the digital image processing process. b. optical character recognition (ocr) ocr takes care of the problem of recognizing optically processed characters. optical recognition is done offline as well as online. offline after writing or printing is complete whereas online recognition is done where the computer recognizes the characters as they are drawn. printed and/or handwritten characters are recognizable but the results directly depend on the quality of the input document. the more limited the input, the better the performance of the ocr system. but when it comes to the completely unrestricted handwriting performance of the ocr engine it is questionable. figure 2.3 shows a schematic representation of the various character recognition areas [2]. figure 2 schematic of optical character recognition (ocr.) area) c. template matching correlation template matching is a technique in digital image processing that has a function to match each part of an image with the image that becomes the template/reference [4]. this is done by comparing the input image with the template image in the database, then looking for similarities using a certain rule. the image matching process that produces a high level of similarity / similarity determines that an image is recognized as one of the template images. d. kiln manufacture furnace or also often referred to as a combustion furnace is a device used for heating. the name comes from the latin fornax, oven. sometimes people also call it a kiln. a kiln is a tool or installation designed as a place of combustion using certain fuels that can be used to heat something [1]. the furnace is simple, composed of stones arranged so that the fuel is protected and heat can be directed. in manufacturing companies, the furnace is made in such a way that the fire or heat that is formed is not too dangerous for the user. klin at pt xyz, the kiln used is a single layer tunnel kiln which consists of 6 combustion zones, namely sub dryer, pre heating, firing, rapid cooling and cooling, with asbestos insulation and using lng (liquid natural gas) as fuel. the fuel will produce heat energy for the ceramic tile burning process. the heat that has been used in the firing section is not completely removed. most of it will be reused to flow to the dryer and sub dryer. but there is also some heat energy that is wasted because it contains carbon which can affect the results of tile products. image acquisition (image capture) image quality improvement image representative process 55 | vol.3 no.2, july 2022 2.2 research method this research is applied research to identify character in image in kiln machine by using optical character recognition (ocr) method. based on the identification of problems obtained in the field observation process, literature study and interviews, namely during the combustion process in the kiln engine, so an application system was created to convert image data into alphanumeric data, using the optical character recognition (ocr) method based on template matching, which then results the conversion can be saved into a database. the research method can be seen in figure 3. figure 3 research flowchart block 2.3 prototype architectural design figure 4 depicts the prototype architecture for monitoring the parameters of the ceramic tile combustion engine (kiln manufacture) for intensive monitoring of the engine. figure 4 prototype architecture the prototype of this ocr data processing application was made by following the steps shown in figure 5 below: figure 5. system workflow 2.4 test design the result of the trial on the application is the result parameter in the conversion of image data into alphanumeric data using ocr. the resulting parameters will be used as a basis for predicting the value that will come out. in order to know how accurate, the prediction will be, before the parameter is formed, the parameter is first evaluated and validated by using a confusion matrix calculation consisting of accuracy, precision and recal [5]. the confusion matrix table can be seen in table 1. tabel 1 tabel confusion matrix true value true false predicted value true tp fp (true positive) (false positive) corect result unexpected result false fn tn (false negative) (true negative) missing result corect absence of result so, the precision, recall and accuracy formulas can be seen in formulas 1, 2 and 3. precision = 𝑇𝑃 𝑇𝑃+𝐹𝑃 1 recall = 𝑇𝑃 𝑇𝑃+𝐹𝑁 2 accuracy = 𝑇𝑃+𝑇𝑁 𝑇𝑃+𝑇𝑁+𝐹𝑃+𝐹𝑁 3 image identification preprocessing result identification developing application models prototyping dokumantatio capture image segmentation grayscaling and binaryization optical character recognition extracted text 56 | vol.3 no.2, july 2022 iii. results and discussion 3.1 preparation of training data and test data in this study, as many as 100 image data taken from the kiln machine in .png format which will later be prepared for training data and test data which is the result of web capture which will later be converted into alphanumeric data. the figure contains parameters that must be monitored intensively for 24 hours. these parameters contain the temperature, gas pressure and air humidity which are read by sensors in certain parts along the ceramic tile combustion engine. at this writing, 10 of the 39 parameters were taken as samples for testing the conversion of images to alphanumeric data using ocr based on template matching. in detail can be seen in figure 6 table 2 below. figure 6 capture images of the kiln machine. tabel 2 description sample parameters for the ocr conversion process. no parameter description 1 tr1 preheating zone, which is the initial zone or area for the product to be heated before being burned in the next zone. the unit used is degrees celsius (ºc). 2 tr3 is firing zone no.1 compaction process (pressure) at high temperatures so that changes in microstructure occur. in this parameter the unit used is degrees celsius (ºc). 3 tc2 it is a thermocontroller parameter no2 to measure the temperature in the combustion process in the zone before firing zone no 3. the unit used is degrees celsius (ºc). 4 tc3 in this zone, thermocontroller parameter no 3 is used to measure the temperature in the combustion process in the zone before firing zone no 4. the unit used is degrees celsius (ºc). 5 tr14 is the zone after passing through the cooling zone no. 3 or to lower the temperature before entering the sub dryer. the unit used is degrees celsius (ºc). 6 tr17 it is a drying area (sub dryer) after cooling the product. the parameter unit used is degrees celsius (ºc). 7 kch is a parameter to see the hydraulic pressure (hydraulic kiln car) in running the conveyor while the machine is running. the unit used is kg/cm² 8 ti3 is the area to measure the temperature of the product no. 3 (zone temperature) before the product comes out of the machine after the drying process. the parameter unit used is degrees celsius (ºc). 9 ti7 is an area to measure the product temperature no. 7 (zone temperature) before the product comes out of the machine after the drying process. the parameter unit used is degrees celsius (ºc). 10 hi3 is a sensor to measure the humidity of the air in the kiln machine in zone no 3 (zone humidity) 3.2 modelling making a model or prototype at this writing, the author makes a prototype which is divided into 3 parts, the first for processing ocr data using thinker board from asus, the second dummy or imitation to display the image of the kiln machine using oracle vm virtualbox, and the third prototype for the application server. the ocr uses oracle vm virtualbox to represent it as a documentation server. 3.3 character recognition the ocr system was created using the tesseract ocr software which was run on python 3.7 for ocr recognition from converted and segmented binary images. tesseract ocr will recognize each character from the segmentation results in the image after previously training the character template. the process of recognizing the character of the kiln machine image on the template uses four parameters when initialized, namely data path, language, mode, and white list [6] so that to obtain accurate detection results, a template is created as a path source as shown in table 4.3. below this: 57 | vol.3 no.2, july 2022 table 3 implementation of the tesseract ocr template no citra hasil segmentasi template karakter 1 0 2 1 3 2 4 3 5 4 6 5 7 6 8 7 9 8 10 9 11 . 3.4 ocr application research results the results of the research on the prototype model of data conversion using ocr on the kiln machine pt. xyz can be seen from table 4 for type a tile products and table 5 for maroon 1000918 products below: tabel 4 ocr conversion result for type a tile products tabel 6 ocr conversion result for product maroon100918 3.5 ocr implementation test results the test results of the ocr conversion prototype model on the kiln machine that have been made using 11 training data as in table 4.2 and 100 test data containing 1000 parameters which are divided into two types of tasks / products can be seen in table 4.4 as many as 520 parameters for type a tile products and in table 4.5 there are 480 parameters for the maroon 100918 tile product. 58 | vol.3 no.2, july 2022 the results of testing the application of ocr can be seen in table 7 below. table 7 confusion matrix test results result class current result class current current 520 0 result 0 480 from table 5 and table 6 above, both the parameters of the type a roof tiles and maroon100918 products obtained accurate conversion results and it is certain that there are no inappropriate parameters. from table 7, the following accuracy is obtained: accuracy = 𝑇𝑃+𝑇𝑁 𝑇𝑃+𝑇𝑁+𝐹𝑃+𝐹𝑁 = 520+480 520+480+0+0 = 1000 1000 = 1 x 100% = 100 % based on the results of these calculations, the accuracy of 100.0% is obtained. iv.conclusion based on the discussion of the results of the research and testing of the research above, it can be concluded that the application model for converting images to numerical data using the optical character recognition (ocr) method based on template matching obtained an accuracy of 100.00%. references [1] arnofiandi, m. s. (2020) ‘rancang bangun tungku pemanas dalam proses metalurgi serbuk’, pembelajaran olah vokal di prodi seni pertunjukan universitas tanjungpura pontianak, 28(2), pp. [2] chaudhuri, a., mandaviya, k., badelia, p., and ghosh, soumya k. (2017) optical character recognition systems, studies in fuzziness and soft computing. doi: 10.1007/978-3-319-50252-6_2 [3] putra, d. (2010) ‘pengolahan citra digital’. yogyakarta: c.v andi offset (penerbit andi), p. 420. [4] hossain, m. a. and afrin, s. (2019) ‘optical character recognition based on template matching’, global journal of computer science and technology, 19(2), pp. 31–35. doi: 10.34257/gjcstcvol19is2pg31 [5] andono, p. n., sutojo, t. and muljono (2017) pengolahan citra digital. yogyakarta: penerbit andi. [6] nunamaker, b. bukhari, s., borth, d. and dengel, a. (2016) ‘a tesseract-based ocr framework for historical documents lacking ground-truth text german research center for artifical intelligence (dfki) kaiserslautern , germany univeristy of kaiserslautern , gemany’, ieee international conference on image processing (icip), pp. 3269–3273 p-issn : 2715-2448 | e-issn : 2715-7199 vol.4 no.1 january 2023 buana information technology and computer sciences (bit and cs) 24 | vol.4 no.1, january 2023 cardiovascular disease prediction using machine learning shivampandey undergraduate student of be cse aiml, chandigarh university, punjab email: shivampandey3819@gmail.com ‹β› abstract—because of technology developments, the ecg yields improved outcomes in the realm of biomedical science and research. the electrocardiogram reveals basic the heart's electrical activity. early detection of aberrant heart disorders is crucial for diagnosing cardiac problems and averting sudden cardiac deaths. measurements on an electrocardiogram (ecg) among people with comparable cardiac issues are essentially equal. analyzing the electrocardiogram characteristics can help predict abnormalities. medical professionals presently base the preponderance of their electrocardiogram diagnosis on their unique particular areas of expertise, which places a substantial load on their shoulders and reduces their performance. the use of technology that automatically analyses ecgs as hospital personnel performs their duties will be advantageous. a suitable algorithm must be able to categories input signal with uncertain awesome feature on just how much they approximate input signal having known characteristics in order to speed up the identification of heart illnesses. a possibility of identifying a tachycardia is raised if this predictor can reliably recognize connections, and this technique may be helpful in lab settings. to accurately diagnose myocardial illness, a powerful machine learning technique should be used. through using recommended method, the effectiveness of cardiovascular disease identification using ecg dataset was evaluated. the reliability, sensitivities, and validity obtained using the svm algorithm were 99.314%, 97.60%, and 97.60% respectively. keywords— machine learning, heart disease, cardiovascular , dataset, engineering i. introduction though electrophysiological (ecg) readings are records of the bioelectrical activity of the cardiovascular system, when cardiovascular disease manifests, most or all of the indications diverge beyond their steady levels [1]. ecg measurements from individuals who have comparable heart problems are quite comparable. although the parameters of an electrocardiogram are unavailable, it is possible to assumed that the signal has the same arrhythmia if its structural pattern mimics that of an electrocardiogram with a particular irregularity. the ecg signal pattern can occasionally be analyzed to recognize heart problems. one of the most difficult and significant health challenges in the real world is coronary heart diagnosis. such condition has an influence on the blood vessel functioning, which might weaken the person's body. in accordance with the who, around 18 million people die from cardiovascular disease each and every year. because heart diseases are becoming increasingly common, individuals are more likely to avoid fatal situations. these are used to assess a patient's level of cardiovascular health [2]. wearable sensors can be used to detect diseases like cardiovascular disease, but they are susceptible to failure because of signal abnormalities. the validity of the data and the test results may be impacted by this problem. in order to anticipate and evaluate a variety of cardiovascular disorders, data mining and hybrid models have been suggested as prospective remedies. with textual information, different risk characteristics are extracted using a data mining approach. in ml algorithms, there still are essentially different phases. the very first permits for the selection of a trait's subset or importance, whilst the second forecasts the development of cardiovascular disease. inserting unnecessary characteristics may generate disturbance and misunderstanding whenever creating a base classifier. further, managing individuals might make categorization less accurate. it's indeed normal practice to discriminate among characteristics and classifications using uncertain combina tions processes. variables might increase the mean square error and lower your accurateness. an ecg is a non-invasive diagnostic tool used to track the physiological activities of the cardiac. it is capable of detecting a variety of circulatory disturbances, including those brought on by myocardial injury, cardiac arrhythmia, and acute coronary syndrome (sca). it is difficult to regularly examine wearable ecg monitoring because of their quick advancement and widespread affordability. machine learning algorithms methods have been extensively applied in a variety of fields, such as video processing, computational linguistics, and automatic speech recognition. their order to transfer extracted features but without help of user experts has been one of many main advantages. alternatively, operations were carried out by algorithms using its capacity for data-driven learning. timely screening of irregular heartbeat circumstances is crucial for the management of unexpected cardiac death and some other acute ailments carried on by myocardial infarction. a number of studies conducted on examining electrocardiogram information and identifying problems therein. 25 | vol.4 no.1, january 2023 to find these anomalies inside this lab, each person wants to just have continued electrocardiogram measurement. this procedure takes a lot of time and energy. automating the automatic data treatment based on computer software is a quicker tool for diagnosing heart problems. an innovative computerized segmentation method is discussed in this work that can classify comparable ecg signals among distinct categories and predict the likelihood of heart disease in each classification. ii. literature review machine learning in field of medical machine learning (ml), a branch of ai technology, offers a selection on novel techniques and methodologies for developing statistics interpretive and forecasting predictions. doctors make a disease diagnosis based on their training, knowledge, observational studies, and expertise. machine learning might prove to be extremely useful in helping people acquire more about and comprehend healthcare. utilizing machine learning (ml) systems for precise diseases diagnosis and prevention in accordance with clinical indications and feelings, healthcare experts have correctly diagnosed sufferers [9],[10],[11],[12]. utilizing electrocardio graphic information, clinical characteristics, and intelligent systems, detect, categories, evaluate the seriousness of, and prognostic adverse reactions in cardiovascular. a unique video-based computational intelligence method called echo net-dynamic was successfully produced by [5]. in comparison to human analysts, our algorithm can evaluate echocardiography footage to estimate cardiovascular system [3]. in biomedical sciences, the physician will be able to diagnose the patient plus figure out the best course of treatment with the help of the pulse rate, electrocardiogram, and blood oxygen data that have been acquired. actual analyzing techniques and internet of things technology can assist warn sufferers concerning impending cardiovascular catastrophes [4],[13],[14],[15]. figure 1. ml method for cardiovascular disease diagnosis. several steps needed to build a model of machine learning for prediction and diagnosis are shown in figure 1. the first phase is gathering pertinent clinical evidence. after becoming sanitized, the data is split into two sets: training and assessment. svm, lr, k-nn, and other computer vision (ml) methods are used to build the model using training examples. the performance and reliability of the model are evaluated using the testing data. the last step could be to either choose a totally opposite model or enhance the efficiency of the current model that includes additional characteristics. iii. method a. supervised machine learning this was used to also create a forecasting model, which forecasts the upcoming based on the historical data. this teaching algorithm employs intake of classification model to complete the task on time. forecasting and classifications tasks belong to the category of supervised methods. using past precipitation data, for example, to forecast rainfall (regression task). by using photographs of salmon, the with tag "fish," the algorithm is expected to identify squid pictures and finish the multiclass classification [5]. b. algorithm used decision tree: although it could be used to handle machine learning problems, the supervised machine learning method termed as a tree structure was mostly usually utilized to resolve detection problem. this technique basically divides the entire set-in smaller chunks while constructing a tree visualization of the data, where every other node in the tree standing in for a classification model and indeed the interior nodes represent judgment nodes, and reflect the attributes. this characteristic at every cluster that separates the classification model most effectively is selected by the procedure. figure 2. showing working of decision tree as example, this clustering algorithm in fig.2 predicts whether such a person will probably purchase a computer. the tree is generated using training data, or tuples with known class labels. when a individual is a learner, bifurcation is based upon their age and credit history. the attribute values for a certain new tuple were contrasted with the tree structure. a route that shows the class predictions for the combination is constructed first from base to a binary tree [6]. c. the fundamental algorithm steps the step is performed sequentially and highest. every one of the training datasets are situated right at the bottom of the tree. categorical values are recommended for attribute values. before being employed in the model, continuous values are discretized. 26 | vol.4 no.1, january 2023 figure 3. showing barplot of the dataset divisions of training sample are iteratively constructed based on the given parameters. a quantitative calculation is used to choose the diverging properties (such as information benefit, gain ratio, or gini index). the splitting cycle is continued until every occurrence of a specific node is a member from same category. probably there are no other input characteristics or there are not enough observations left to provide an accurate split. evaluate the strategy with data, then determine whether it is accurate. d. support vector machine typically used only for classifications, svms are indeed a popular category of supervised algorithms for machine learning. the svm classifier converts characteristic data into coordinates in an n-dimensional area. the information is then classified by a higher dimensional space that the program finds. the classifier represents a maximum margin. the fundamental idea behind svm is to repeatedly find a greatest margin class label (mmh) that accurately categories the collection only with fewest errors [7]. e. how does svm algorithm work promote efficient hyperplane that separates to divide the categories. picture just on left showing various black, blue, and orange optimal hyperplane. even though the black in this scenario adequately distinguishes the 2 classes, the blue and orange exhibit greater classifying errors. hyperplane: a hyper plane is a judgement layer that makes distinctions between such a collection of elements with varied class affiliations. margin is the separation here between two on the closest classification points. an angle between both the line and the nearby points or testing set is used to determine value. a massive class difference is seen to be beneficial; a reduced class difference is considered to be harmful. the descriptor with the highest separation as from closest point should be chosen. iv. results and discussion in the purpose of our research would have been to concurrently mitigate underfit and overfit defects in some other good design. we found and observed the model didn't result either in the overfitting or underfitting. it is expected that the system loss in training data will be less than that in testing data. another benefit is that if we understand these crucial ideas and are aware of how effective, we can handle even the most stressful circumstances. in comparison to test data, prototype loss should have been lower in training instances. there were two divisions used to rate the accuracy. our system improved as the quantity of training photos and parameter settings increased, resulting in. v. conclusion throughout the healthcare profession, cardiovascular disease diagnosis is difficult and crucial. the detection of cardiovascular problems through to the examination of unprocessed medical data will aid inside the lengthy safeguarding of human life. if indeed the illness is identified in its beginning phases and protective actions were implemented as quickly as feasible, the number of deaths could be managed. this aids with in earliest diagnosis of cardiovascular problems. therefore, in study, the svm classifier is used to gather data and provide a strategy for predicting cardiovascular disease with a reliability of 97.60%. to concentrate the researches on actual data rather than conceptual techniques and computations, a further development of something like the research is highly necessary. references [1] p. mcsharry, g. clifford, l. tarassenko, method for generating an artificial rrtachogram of a typical healthy human over 24-hours, comput. cardiol. 29(2002) 225–228. [2] s. jayalalitha, d. susan, shalini kumari and b. archana, “k-nearest neighbour method of analysing the ecg signal (to find out the different disorders related to heart)”, journal of applied sciences, 14: 1628-1632 [3] romiti s, vinciguerra m, saade w, anso cortajarena i, greco e. artificial intelligence (ai) and cardiovascular diseases: an unexpected alliance. cardiol res pract. 2020 jun 27;2020:4972346. [4] mamun, m.m.r.k. significance of features from biomedical signals in heart health monitoring. biomed 2022, 2, 391-408. [5] energy fuels 2022, 36, 13, 6626–6658 publication date:june 13, 2022 [6] rokach, lior & maimon, oded. (2005). decision trees. 10.1007/0-387-25465-x [7] han, j., and m. kamber. 2011. data mining: concepts and techniques. 3rd ed. burlington: morgan kaufmann. [8] joachims, t. 1998. making large-scale svm learning practical. adv. kernel methods support vector learn, mit press. [9] a. m. shah et al., “echocardiographic features of patients with heart failure and preserved left ventricular ejection fraction,” j. am. coll. cardiol., vol. 74, no. 23, pp. 2858–2873, 2019. [10] s. horiuchi and j. p. kneller, “what can be learned from a future supernova neutrino detection?,” j. phys. 27 | vol.4 no.1, january 2023 g nucl. part. phys., vol. 45, no. 4, p. 43002, 2018. [11] m. a. lancaster and m. huch, “disease modelling in human organoids,” dis. model. mech., vol. 12, no. 7, p. dmm039347, 2019. [12] s. j. al’aref et al., “clinical applications of machine learning in cardiovascular disease and its relevance to cardiac imaging,” eur. heart j., vol. 40, no. 24, pp. 1975–1986, 2019. [13] j. yu, w. ouyang, m. l. k. chua, and c. xie, “sarscov-2 transmission in patients with cancer at a tertiary care hospital in wuhan, china,” jama oncol., vol. 6, no. 7, pp. 1108–1110, 2020. [14] j. stehlik et al., “continuous wearable monitoring analytics predict heart failure hospitalization: the linkhf multicenter study,” circ. hear. fail., vol. 13, no. 3, p. e006513, 2020. [15] a. a. kulkarni, v. e. vijaykumar, s. k. natarajan, s. sengupta, and v. s. sabbisetti, “sustained inhibition of cmet-vegfr2 signaling using liposome-mediated delivery increases efficacy and reduces toxicity in kidney cancer,” nanomedicine nanotechnology, biol. med., vol. 12, no. 7, pp. 1853–1861, 2016. p-issn : 2715-2448 | e-isssn : 2715-7199 vol.3 no.1 january 2022 buana information tchnology and computer sciences (bit and cs) 1 | vol.3 no.1, january 2022 sha512 and md5 algorithm vulnerability testing using common vulnerability scoring system (cvss) fahmi basya 1 study program information technology faculty universitas budi luhur 1911600482@student.budiluhur.ac.id mardi hardjanto 2 study program information technology faculty universitas budi luhur mardi.hardjianto@budiluhur.ac.id ‹β› ikbal permana putra 3 study program information technology faculty universitas budi luhur 2011600364@student.budiluhur.ac.id abstrak—tulisan ini membahas perbandingan hasil pengujian algoritma otp (one time password) pada dua enkripsi yaitu sha512 dan md5 yang diterapkan pada aplikasi rekonsiliasi dinas pemberdayaan masyarakat dan desa kabupaten sukabumi. studi ini menggunakan metode vulnerability assessment and penetration testing (vapt), yang menggabungkan dua bentuk pengujian kerentanan untuk mencapai analisis kerentanan yang jauh lebih lengkap dengan melakukan tugas yang berbeda di area fokus yang sama. penilaian kerentanan menggunakan metode common vulnerability scoring system (cvss). hasil penelitian menunjukkan bahwa metode vulnerability assessment and penetration testing (vapt) terbukti mampu mengidentifikasi tingkat kerentanan keamanan pada aplikasi rekonsiliasi di dinas pemberdayaan masyarakat dan desa kabupaten sukabumi dengan skor tingkat kerentanan 5,3 pada lingkungan sha512 dengan peringkat sedang dan 7,5 di lingkungan md5. dengan peringkat tinggi. jadi, dapat disimpulkan bahwa algoritma terbaik untuk mengimplementasikan otp adalah sha512. kata kunci— otp, sha512, md5, vapt, cvss abstract—this paper discusses the comparison of the results of testing the otp (one time password) algorithm on two encryptions, namely sha512 and md5 which are applied to the reconciliation application of the dinas pemberdayaan masyarakat dan desa kabupaten sukabumi. this study uses the vulnerability assessment and penetration testing (vapt) method, which combines two forms of vulnerability testing to achieve a much more complete vulnerability analysis by performing different tasks in the same focus area. the vulnerability assessment uses the common vulnerability scoring system (cvss) method. the results showed that the vulnerability assessment and penetration testing (vapt) method was proven to be able to identify the level of security vulnerability in the reconciliation application at the dinas pemberdayaan masyarakat dan desa kabupaten sukabumi with a vulnerability level score of 5.3 in the sha512 environment with a medium rating and 7.5 in the md5 environment. with high ratings. so, it can be concluded that the best algorithm for implementing otp is sha512. keywords— otp, sha512, md5, vapt, cvss i. introduction the dinas pemberdayaan masyarakat dan desa kabupaten sukabumi is a government agency that applies information technology to support daily work processes. the reconciliation application which is one of the assets of the dinas pemberdayaan masyarakat dan desa kabupaten sukabumi is a village financial application that functions to control village finances which includes financial input and financial realization. however, at this time, the reconciliation application has not implemented a good security system, so application security is currently very vulnerable to cyber-attacks. in addition, to ensure its security, it is necessary to test the system security on the application. observing this problem, it is necessary to improve system security by implementing two factor authentication, one of which uses one time password (otp). otp is an authentication method that uses a password that always changes every login, or changes every certain time interval. it can also be called a password that is only valid for a single login session [1]. the otp algorithms tested in this study were sha512 and md5. sha512 is an algorithm that uses a one-way hash function created by ron rivest [2] and is a development of the sha0, sha1, sha256, and sha384 algorithms. the hash itself functions to accept an input string of any length and convert a string whose output length remains the same [3]. the sha512 algorithm itself is used when generating random codes as an otp code generator. the sha512 function produces a message digest with a size of 512 bits and a block length of 1024 bits. in sha512 there are 80 rounds and for padding bits it is done as in sha-1, but the block size in sha512 becomes 1024 bits [4]. while the md5 algorithm was designed by ron rivest whose use is very popular among the open-source community as a checksum for downloadable 2 | vol.3 no.1, january 2022 files [5] which has a block size of 512 bits with a digest size of 128 bits. the web security testing tool for both sha512 and md5 algorithms in this study is the burp suite, which is a java-based integrated platform for conducting security testing of web applications. burp suite in general is a web penetration testing framework. this tool is specially made for web applications. burp suite is now used by most professional testers as part of industry standard penetration tools. in the system security testing process, a combination of two forms of vulnerability testing is used to achieve a much more complete vulnerability analysis by performing different tasks in the same focus area, known as vulnerability assessment penetration testing (vapt) [6]. therefore, in this study, vulnerability testing of the reconciliation application system will be carried out using the vapt method [9]. ii. method the research method used is the penetration testing method which refers to a security test framework for a web application system, namely vulnerability assessment penetration testing (vapt). vapt itself is a combination of two vulnerability assessment and penetration testing activities, where vulnerability assessment is an activity that includes the process of examining a security vulnerability of a web application. while penetration testing is a process of simulating attacks on vulnerabilities found on the web and exploring them. the testing process using the vatp framework contains 9 necessary stages as shown in the following figure [7]. figure 1. vapt framework based on figure 1, this testing phase begins with the scope step (determining the focus of the test, namely information technology assets belonging to the dinas pemberdayaan masyarakat dan desa kabupaten sukabumi, namely the reconciliation application), reconnaissance (at this stage information is collected as a basis for targets related to the system will be tested for vulnerability , such as ip address and hosting using domain tools), vulnerability detection (at this stage the process of finding information about security vulnerabilities in the reconciliation application using the burp suite tools on the otp that has been applied is carried out. after finding vulnerabilities, the results are then used as a basis in the planning stage next), information analysis and planning (at this stage analysis and test planning is carried out, namely looking for security vulnerabilities. this analysis and planning will later be used in the penetration testing process), penetration testing (at this stage sim ulation of attacks on information technology assets of the dinas pemberdayaan masyarakat dan desa kabupaten sukabumi, namely the reconciliation application, namely by brute force), privilege escalation (at this stage the vulnerability exploitation process is carried out by utilizing information on vulnerabilities from the results of the penetration testing process), result analysis, reporting and clean-up (at this stage the preparation of a report on the results of the previous stages of testing is carried out). vulnerability assessment refers to the standardization of owasp (open web application security project) as well as penetration testing by conducting penetration testing according to standards. this study then compares the results of the application of the two sha512 and md5 algorithms in the reconciliation application of the dinas pemberdayaan masyarakat dan desa kabupaten sukabumi through the otp flow in the figure 2. based on figure 2, the flow of the application that will be developed using sha512 and md5 to generate a time-based otp code is synchronized directly with the server in the verification of the otp code so that it can access into the system. the steps for generating the otp code are that the user enters the username and password on the login page, then it will be processed by sha512 or md5, then the otp code will be sent to the user's telegram application who will log in [8]. furthermore, the two algorithms are tested and assessed using the cvss calculation in table 1 and figure 3 below. figure 2. otp flow table 1. cvss score rating base score none 0,0 low 0,1 – 3,9 medium 4,0 – 6,9 high 7,0 – 8,9 critical 9,0 – 10,0 clean-up reporting result analysis privilege escalation penetration testing information analysis ang planning vulnerability detection reconnaissance scope 3 | vol.3 no.1, january 2022 figure 3. cvss calculation base score: if (impact sub score <= 0) else, 0) scope unchanged round up (minimum [(impact + exploitability), 10]) scope changed round up (minimum [1.08 x (impact + exploitability),10]) impact sub score scope unchanged 6.42 x iscbase scope changed 7.52 x [iscbase-0.029]-3.25 x [iscbase0.02] impact sub score: iscbase = 1 – [(1-impactconf) x (1-impactinteg) x (1impactavail)] exploitability sub score: 8.22 x attackvector x attackcomplexity x priviligerequired x userinteraction vulnerability level testing uses a scoring system from the (cvss) common vulnerability scoring system as a standard for calculating the level of a vulnerability in the system with a vulnerability level value as shown in table 2. table 2. vulnerability level of cvss rating base score range none 0,0 low 0,1 – 3,9 medium 4,0 – 6,9 high 7,0 – 8,9 critical 9,0 – 10,0 iii. results and discussion the level scoring test refers to the common vulnerability and exposure (cve) where cve is a list that displays any security information, both on software and firmware that are vulnerable to cyber-attacks. this test is done technically to prove the scoring level on vulnerability. after obtaining vulnerabilities for both the sha512 environment and md5 algorithms, further testing is carried out on the technical login page (intruder positions process). this test is carried out on the sha512 environment and md5 algorithms using tools, namely the burp suite as a scanner as well as the executor, so that the target being analyzed has bugs so that the system can be hacked. testing the vulnerability of this target using brute force techniques. tabel 3. cvss formula sha512 base metric evaluasi skor attack vector network 0,85 attack complexity high 0,44 privileges required low 0,62 user interaction none 0,85 scope unchanged 0,68 confidentiality high 0,56 integrity none 0 availability none 0 base formula cvss:3.1/av:n/ac:h/pr:l/ui:n/s:u/c:h/i:n/a:n iss = 1 [(1-confidentiality) × (1-integrity) × (1-availability)] impact if scope is unchanged = 6.42 × iss exploitability = 8.22 × attackvector × attackcomplexity × privilegesrequired × userinteraction iss = 1 [ (1 – 0.56) × (1 0) × (1 0)] = 1 (0.44) = 0.56 impact if scope is unchanged = 6.42 × 0.56 = 3.6 exploitability = 8.22 × 0.85 x 0.44 x 0.62 x 0.85 = 1.6 impact + exploitability = 3.6 + 1.6 = 5.3 base score if scope is unchanged = round up (5.3, 10) = 5.3 after the intruder positions are executed, the payload sets configuration is then performed, before brute force testing is performed. at this stage, the burp suite will input the possible username and password used. furthermore, brute force attack testing was carried out on the sha512 environment and md5. then testing the results of the brute force attack on the rendering menu, it appears directly on the otp page. furthermore, the configuration of the intruder positions and payload before testing is carried out. the penetration process was successfully carried out by obtaining an otp code when testing with a brute force attack was carried out. the next process equates the otp code sent on telegram with the test results on the sha512 environment and md5. the results of testing the sha512 environment vulnerability level are presented in table 3 and figure 4. figure 4. base score metrics sha512 based on table 3 and figure 4, it is found that the vulnerability level of the sha512 algorithm is at the medium level. the results of testing the md5 vulnerability level can be seen in table 4 and figure 5. 4 | vol.3 no.1, january 2022 tabel 4. cvss formula md5 base metric evaluasi skor attack vector network 0,85 attack complexity high 0,44 privileges required low 0,62 user interaction none 0,85 scope unchanged 0,68 confidentiality high 0,56 integrity none 0,56 availability none 0.56 base formula cvss:3.1/av:n/ac:h/pr:l/ui:n/s:u/c:h/i:n/a:n iss = 1 [ (1 confidentiality) × (1 integrity) × (1 availability)] impact if scope is unchanged = 6,42 × iss exploitability = 8,22 × attackvector × attackcomplexity × privilegesrequired × userinteraction iss = 1 [ (1 – 0,56) × (1 – 0,56) × (1 – 0,56)] = 1 (0,085) = 0,914 impact if scope is unchanged = 6,42 × 0,914 = 5,9 exploitability = 8.22 × 0.85 x 0.44 x 0.62 x 0.85 = 1.6 impact + exploitability = 5,9 + 1,6 = 7,5 base score if scope is unchanged = round up (7,5, 10) = 7,5 figure 5. base score metrics md5 based on figure 4, it is obtained that based on figure 5, it is found that the md5 algorithm vulnerability level is at a high level. iv. conclusion the results of testing security vulnerabilities in the reconciliation application of the dinas pemberdayaan masyarakat dan desa kabupaten sukabumi using the vulnerability assessment and penetration testing (vapt) method and the blackbox testing approach resulted in the level of vulnerability values being medium in the sha512 environment and in the md5 environment with high vulnerability values. from the test results, it is identified that there is a vulnerability in the browser session and the use of an inappropriate hash function, which has the potential to carry out a brute force attack. references [1] perdana, u. p. s. (2016) ‘pemanfaatan telegram bot api dalam layanan otentikasi tanpa password menggunakan algoritma time-based one-time password (totp)’, pp. 1–12. [2] juardi, d. (2017) ‘kajian vulnerability keamanan data dari eksploitasi hash length extension attack vulnerability data satisfaction study from exploitation hash length extension attack’, 6(1) [3] rizki, r. and mulyati, s. (2020) ‘implementasi one time password menggunakan algoritma sha-512 pada aplikasi penagihan hutang pt. xht’, edumatic : jurnal pendidikan informatika, 4(1), pp. 111–120. doi: 10.29408/edumatic.v4i1.2158. [4] sembiring, j. (2013) ‘analisis algoritma sha-512 dan watermarking dengan metode least significant bit pada data citra’, seminar nasional sistem informasi indonesia, pp. 2–4. [5] sulastri, s. and putri, r. d. m. (2018) ‘implementasi enkripsi data secure hash algorithm (sha-256) dan message digest algorithm (md5) pada proses pengamanan kata sandi sistem penjadwalan karyawan’, jurnal teknik elektro, 10(2), pp. 70–74. doi: 10.15294/jte.v10i2.18628 [6] simran, g. and sasikala, d. (2019) ‘vulnerability assessment of web applications using penetration testing’, international journal of recent technology and engineering, 8(4), pp. 1552–1556. doi: 10.35940/ijrte.b2133.118419. [7] goel, j. n. and mehtre, b. m. (2015) ‘vulnerability assessment & penetration testing as a cyber defence technology’, procedia computer science, 57, pp. 710– 715. doi: 10.1016/j.procs.2015.07.458. [8] setiawan, d. a. et al. (2018) ‘implementasi one time password menggunakan algoritma hash sha-512 berbasis web pada badan kepegawaian dan pengembangan sdm kota’, skanika volume 1 no. 1 maret 2018 implementasi, 1(1), pp. 199–204. [9] d. kurniawan, a. l. hananto, and b. priyatna, “modification application of key metrics 13x13 cryptographic algorithm playfair cipher and combination with linear feedback shift register (lfsr) on data security based on mobile android,” int. j. comput. tech.-–, vol. 5, no. 1, pp. 65–70, 2018. [10] b. huda, “sistem informasi data penduduk berbasis android dan web monitoring studi kasus pemerintah kota karawang (penelitian dilakukan di kab. karawang),” buana ilmu, vol. 3, no. 1, pp. 62–69, 2018, doi: 10.36805/bi.v3i1.456. p-issn : 2715-2448 | e-issn : 2715-7199 vol.4 no.1 january 2023 buana information technology and computer sciences (bit and cs) 1 | vol. 4 no.1, january 2023 coastal batik motifs identification using k-nearest neighbor based on the grey level co-occurrence method wresti andriani1 informatics engineering stmik ymi tegal email: wresty.andriani@gmail.com gunawan2* informatics engineering stmik ymi tegal email: gunawan.gayo@gmail.com sawaviyya anandianskha3 informatics engineering stmik ymi tegal email: sayaviyyaa@gmail.com ‹β› abstrak-indonesia merupakan negara yang kaya akan sumber daya alam, budaya dan pariwisata. salah satu warisan budaya manusia yang terkenal di indonesia adalah batik. batik memiliki keunikan motif yang sangat beragam sehingga sulit untuk mengenali golongan tertentu terutama generasi muda. penelitian ini dilakukan untuk mengklasifikasikan batik pesisir, khususnya batik tegal, batik pekalongan, dan batik cirebon sehingga dapat membantu memudahkan pengenalan dan pemahaman batik pesisir jika dibandingkan dengan batik pedalaman, seperti batik yogyakarta. metode yang digunakan adalah gray level co-occurrence matrix (glcm) untuk mengekstrak fitur tekstur, sedangkan untuk menentukan kedekatan citra uji dengan data latih menggunakan metode k-nearest neighbor (knn), perhitungan jarak yang digunakan adalah euclidean distance dan manhattan distance berdasarkan karakteristik tekstur dari citra batik yang diperoleh. hasil yang diperoleh pada penelitian ini dimana skor tertinggi adalah 64% untuk euclidean distance dan 66% untuk manhattan distance pada k = 15. kata kunci— pesisir, glcm, knn, identifikasi abstract-indonesia is a country rich in natural resources, culture, and tourism. one of the famous human cultural heritage in indonesia is batik. batik has unique motifs that are very diverse, making it difficult to recognize certain groups, especially the younger generation. this research was conducted to classify coastal batik, especially tegal batik, pekalongan batik, and cirebon batik so that it can help facilitate the introduction and understanding of coastal batik when compared to inland batiks, such as yogyakarta batik. the method used is the grey level co-occurrence matrix (glcm) to extract texture features, while, to determine the proximity of the test image to the training data using the k-nearest neighbor (knn) method, the calculation of the distance to be used is the euclidean distance and manhattan distance based on the texture characteristics of the batik image obtained. the result obtained in this study where the highest score of 64% for euclidean distance and 66% for manhattan distance at k=15. keywords— coastal, glcm, knn, identification i. introduction indonesia is a country rich in culture and beautiful nature. one of the wealth that is owned is the batik culture. indonesian batik varies according to the many cultures that are owned in each region in indonesia. batik is an art that produces picture cloth which is the original heritage of the indonesian nation and is one of the world heritages and has been inaugurated by unesco, the united nations world agency in the fields of culture and education. based on the region, batik can be divided into two types, namely coastal batik and inland batik and based on the manufacturing process, batik is also divided into two, namely written batik, namely batik made by writing use liquid "wax" on cloth and stamp batik, namely batik made by depicting cloth using a stamp that has a pictorial part and is shaped like a relief and inscribed using wax liquid, then processed in a certain way. each batik from various regions has its own uniqueness and high traditional artistic value. the diversity of traditional batik motifs is due to differences in geography, flora and fauna, differences in lifestyle and livelihoods. in the past, batik work was often done by women and was an exclusive job, because the batik cloth produced was presented to the nobility or distinguished guests. coastal batik motives are batik motives produced in coastal areas such as the tegal coast, pekalongan coast, and cirebon coast, while the motives can vary, usually in addition to being influenced by geography and native culture, but also often influenced by culture brought by fishermen or traders from outside the area, making coastal fabric motives more diverse compared to batik motifs from the interior, such as batik motives from yogyakarta which are more influenced by the sense of nobility of the palace. the wide variety of traditional batik motifs in indonesia confuses traders and the younger generation (millennials) in recognizing the origin of existing batik fabrics. because there is no center or office that specifically handles and provides information about the history of batik itself. in this paper, the researcher hopes to help facilitate the millennial generation and related parties to become more familiar with motifs so that they can love these traditional products more, especially tegal batik, pekalongan batik, and cirebon batik. this study, will use the gray level co-occurrent matric (glcm) method as its feature extraction and use the kmailto:wresty.andriani@gmail.com mailto:gunawan.gayo@gmail.com mailto:sayaviyyaa@gmail.com 2 | vol. 4 no.1, january 2023 nearest neighbor (knn) method. the distance calculation uses euclidean distance and manhattan distance to determine the proximity of the test image to the training data and is expected to help identify and classify the image of batik motives. the well-identified image of batik will provide clear information and can be used for the preservation of indonesian batik fabric motifs from extinction. similar studies that have been conducted by some researchers using batik objects include zulfrianto y. lamasigi [1] dct for extraction of glcm-based features on batik identification using k-nn. the highest accuracy obtained by dct-glcm exists at an angle of 135° with a value of k=3 of 64.88% and a value of 64.88% and at an angle of 0° with values k=7 and 9 are 41.86%. frisnanda aditya, etc [2] pekalongan batik identification using the grey level cooccurrence matrix and probabilistic neutral network method. from the results of this classification test, the best accuracy is 61.33%. there is has been no research that examines the identification of the difference between coastal and inland batik using glcm and knn method before. ii. method a. gray level co-occurrence matrix (glcm) grey level co-occurrence matrix (glcm)was first submitted by haralick in 1979 with 28 features to explain spatial patterns. there are several steps that are taken, namely: first, calculating the features of the glcm by converting an rgb image into a grey scale image. second, creating a co-occurrence matrix is continued by determining the spatial relationship between reference pixels and neighbouring pixels based on angle 𝜃 and distance d. third, creating a symmetrical matrix by adding a co-occurrence matrix to the transpose matrix. fourth, the symmetrical matrix is normalized by calculating the probability of each element of the matrix. fifth, calculate the features of the glcm. each feature is calculated by one pixel distance in four directions, i.e. 00, 450, 900, and 1350 to detect cooccurrence [1] [2] [3] [4] [5] [6] [7] [8] [9] [10] [11] [12]. there are 5 glcm features used in this study, including: 1) angular second moment (asm) asm is a measure of the homogeneity of the image. 𝐴𝑆𝑀 = ∑ ∑ (𝐺𝐿𝐶𝑀 (𝑖, 𝑗))2𝐿 𝑗=1 𝐿 𝑖=1 (1) 2) contrast contrast is a measure of the presence of variations in the grey level of images. 𝐶𝑜𝑛𝑡𝑟𝑎𝑠𝑡 = ∑ ∑ |𝑖 − 𝑗|2𝐺𝐿𝐶𝑀 (𝑖𝑗)𝐿 𝑗 𝐿 𝑖 (2) 3) inverse different moment (idm) used to measure homogeneity 𝐼𝐷𝑀 = ∑ ∑ (𝐺𝐿𝐶𝑀(𝑖,𝑗))2 1+(𝑖−𝑗)2 𝐿 𝑗=1 𝐿 𝑖=1 (3) 4) entropy entropy for represents the measure of the grey level irregularity in the image. 𝐸𝑛𝑡𝑟𝑜𝑝𝑖 = − ∑ ∑ (𝐺𝐿𝐶𝑀 (𝑖, 𝑗)) log(𝐺𝐿𝐶𝑀 (𝐼, 𝐽))𝐿 𝑗=1 𝐿 𝑖=1 (4) 5) correlation correlation is a measure of the dependence between the grey values in the image. 𝐶𝑜𝑟𝑟𝑒𝑙𝑎𝑡𝑖𝑜𝑛 = ∑ ∑ (𝑖−𝜇𝑖′)(𝑗−𝜇𝑗′)(𝐺𝐿𝐶𝑀 (𝑖,𝑗)) 𝜎𝑖𝜎𝑗 𝐿 𝑗=1 𝐿 𝑖=1 (5) this equation is based on the mean value of the grey image intensity and the standard deviation. the standard deviation is obtained from the square root of the variant which shows the distribution of pixel values in the image, with the following formula: 𝑚𝑒𝑎𝑛 𝑖 = 𝜇𝑖′ = ∑ ∑ 𝑖 ∗ 𝐺𝐿𝐶𝑀(𝑖, 𝑗) 𝐿 𝑗=1 𝐿 𝑖=1 𝑚𝑒𝑎𝑛 = 𝜇𝑗′ = ∑ ∑ 𝑗 ∗ 𝐺𝐿𝐶𝑀(𝑖, 𝑗) 𝐿 𝑗=1 𝐿 𝑖=1 𝑣𝑎𝑟𝑖𝑎𝑛 𝑖 = 𝜎𝑖2 = ∑ ∑ 𝐺𝐿𝐶𝑀(𝑖, 𝑗)(𝑖 − 𝜇𝑖′)2 𝐿 𝑗=1 𝐿 𝑖=1 𝑣𝑎𝑟𝑖𝑎𝑛 𝑗 = 𝜎𝑗2 = ∑ ∑ 𝐺𝐿𝐶𝑀(𝑖, 𝑗)(𝑗 − 𝜇𝑗′)2 𝐿 𝑗=1 𝐿 𝑖=1 𝑠𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛 𝑖 = 𝜎𝑖 = √𝜎𝑖2 𝑠𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛 𝑗 = 𝜎𝑗 = √𝜎𝑗2 b. k-nearest neighbor k-nearest neighbour (knn) is a method that uses a supervised algorithm the results of the newly classified query instance based on the majority of the categories on the knn. this algorithm aims to classify new objects based on attributes and training samples. the knn algorithm is very simple, based on the shortest distance from the query instance to the training sample to determine its knn. the training sample is projected into a multi-dimensional space, where each dimension represents a feature of the data. space is divided into sections-part based on the classification of the training sample. a point in this space is marked class c if class c is the most common classification in the k nearest neighbour of that point. 1) euclidean distance calculation of the distance of euclidean distance which is represented as follows [3] [13]: 𝑑 = √(𝑎1 − 𝑏1)2 + (𝑎2 − 𝑏2)2 + ⋯ + (𝑎𝑛 − 𝑏𝑛)2 𝑑 = √∑ (𝑎𝑖 − 𝑏𝑖)2𝑛 𝑖=1 (6) where d (a, b): the euclidean distance between the vector a and the vector b, ai: feature vector a, bi: features of the vector b, n: the number of features in the a and b vectors. 2) manhattan distance manhattan or city distance is a similarity measurement that is most suitable for project approvals that represent relevant cases with natural numbers or with quantitative data. also used to retrieve matched cases from the case base by calculating the absolute weighted sum of the differences between the current case and other cases the case base. to calculate the weight, the following quationis used: 3 | vol. 4 no.1, january 2023 𝑑𝑖𝑗 = ∑ 𝑊𝑘 |𝑥𝑖𝑘 − 𝐶𝑗𝑘| (7) 𝑑𝑖𝑗 = 𝑑𝑒𝑠𝑡𝑎𝑛𝑐𝑒 𝑏𝑒𝑡𝑤𝑒𝑒𝑛 𝑐𝑎𝑠𝑒𝑠 𝑖 𝑎𝑛𝑑 𝑗 𝑊 = 𝑟𝑒𝑝𝑟𝑒𝑠𝑒𝑛𝑡 𝑡ℎ𝑒 𝑠𝑢𝑚 𝑜𝑓 𝑤𝑒𝑖𝑔ℎ𝑡 𝑋 = 𝑛𝑒𝑤𝑙𝑦 𝑟𝑒𝑑𝑢𝑐𝑒𝑑 𝑐𝑎𝑠𝑒 𝑤𝑖𝑡ℎ 𝐶 c. confusion matrix the confusion matrix is a table consisting of many rows of data test that predicted true and false by the classification model, to determine the performance of a classification model [14] [15]. table i. confusion matrix predicted class actual class class class = 1 class = 0 class = 1 f 11 f 10 class = 0 f 01 f 00 accuracy calculation using confusion matrix as follows: 𝑎𝑐𝑐𝑢𝑟𝑎𝑐𝑦 = 𝐹11+𝐹00 𝐹 11+𝐹10+𝐹01+𝐹00 or 𝐴𝑐𝑐𝑢𝑟𝑎𝑐𝑦 = 𝑁𝑢𝑚𝑏𝑒𝑟 𝐶𝑙𝑎𝑠𝑠𝑖𝑓𝑖𝑐𝑎𝑡𝑖𝑜𝑛 𝑡𝑟𝑢𝑒 𝑛𝑢𝑚𝑏𝑒𝑟 𝑑𝑎𝑡𝑎 𝑥 100% the flow/stages that will be done in this study can be seen in figure 1 below: figure 1. research flowchart iii. results and discussion data collection (data collection) is obtained from literature study, observation, and data search in the form of batik motive images via the internet. image data obtained are as many as 90 batik motives, consisting of coastal batik from the city of tegal as many as 20 batik motifs, 50 kinds of pekalongan batik motifs, and 20 kinds of cirebon batik motives, also 10 kind of yogyakarta batik motives. batik yogyakarta as representative of inland batik. all the batik data motives obtained are converted into the same size and extension * jpg. this data image is then grouped into training data. yogyakarta batik motif images. all of the test data is the same size and also has a jpg extension. below are some examples of images of batik motives that have been collected: (a) (b) (c) (d) figure 2. example batik (a) from tegal (b) from pekalongan (c) from cirebon (d) from yogyakarta each image is done with the feature extraction process using glcm using matlab 2020. from matlab results, then feature extraction is classified using knn algorithm with 2 kinds of distance measurement, namely euclidean distance and manhattan distance and using rapid miner tool. a. texture feature extraction in the previous process, we obtained a total of 90 data of batik motifs as training data consisting of 20 tegal batik motifs, 50 kinds of pekalongan batik motifs and 20 kinds of cirebon batik motifs along with 10 kinds of batik motifs from yogyakarta as samples of inland batik motifs as a preprocessing stage. the results of this stage are then continued at the extraction stage of the texture feature. the extraction of this feature is done using matlab 2020 which results in table ii below: table ii. glcm results no matrix eccentat ion ... next ... korel ... energy homo… 1 116,570 68,016 0.617 0.005 5,453 0.066 2 90,403 42,895 0.148 0.006 5,134 0.028 3 132,540 56,192 0.210 0.005 5,338 0.046 4 104,627 72,138 0814 0.005 5,318 0.074 5 187,068 46,154 1,568 0.008 5,062 0.032 ⋮ 90 93,465 46,579 0.464 0.006 5,185 0.032 once the glcm feature extraction results using matlab 2020 are obtained, the feature extraction data is obtained from glcm, then the data is entered into the rapid miner tool and 4 | vol. 4 no.1, january 2023 using the loop parameter, it can be determined the value k = 1,3,5,7,9,11,13,15 for used, the k which produces the highest accuracy value to be used from the distance method, the method used are euclidean distance and manhattan (city) distance. the results can be seen in the table 3 below: table iii. comparison between the manhattan and euclidean methods accuracy distance k=1 k=3 k=5 k=7 k=9 k=11 k=13 k=15 manhattan 55 60 61 62 59 61 57 66 euclidean 52 59 60 61 63 62 62 64 the comparison of accuracy in table iii between euclidean distance and manhattan distance above, a graphic image is obtained as shown in figure 3. figure 3. comparison of manhattan and euclidean in the table iii and figure 3, it can be seen that the comparison from euclidean distance and manhattan distance, both of them have the highest accuracy at the value k = 15, where the highest accuracy is obtained at 64% for euclidean distance and is 66% for manhattan distance. the same occurrence for the smallest accuracy value is obtained at the value k=1. manhattan's accuracy is 55% and euclidean’s is 52%. b. confusion matrix in this study, in addition to knowing the comparison between euclidean distance and manhattan city distance, it will also be seen how accurate the coastal batik is with the inland batik by knowing the confusion matrix. the result of the confucius matrix from the classification of coastal batik with inland batik, in this case, represented by tegal batik, pekalongan batik, and cirebon batik for coastal batik and yogyakarta batik as inland batik can be seen from the following table iv. table iv. prediction motive batik performance true cirebon true yogyakarta true pekalongan true tegal class precission pred. cirebon 8 1 3 1 61,54% pred. yogyakarta 0 1 0 4 20,00% pred pekalongan 8 2 44 4 75,86% pred. tegal 4 6 3 11 45,83% class recall 40% 10% 88% 55% 45,83% when the prediction results above are tested using the confusion matrix obtained results in table v below: table v. matrix of the predictions no. label prediction result 1 cirebon pekalongan 0 2 cirebon pekalongan 0 3 yogyakarta pekalongan 0 4 pekalongan pekalongan 1 5 pekalongan pekalongan 1 6 pekalongan pekalongan 1 7 pekalongan pekalongan 1 8 pekalongan pekalongan 1 9 tegal pekalongan 0 10 tegal pekalongan 0 11 cirebon pekalongan 0 12 cirebon pekalongan 0 13 yogyakarta pekalongan 0 14 pekalongan pekalongan 1 15 pekalongan pekalongan 1 16 pekalongan pekalongan 1 17 pekalongan pekalongan 1 18 pekalongan pekalongan 1 the number 0 indicates the wrong label prediction, the number 1 indicates the correct prediction. the accuracy of predictions using manhattan (city) distance is higher than euclidean distance which is 66%. from the batik motif data that has been collected, there is a prediction that the true batik from pekalongan is higher 75.86%, than the true batik from cirebon is 61.54%. it can be concluded data from the training data obtained, after being tested was more predicted as pekalongan batik, then followed by cirebon batik. iv. conclusion based on the results of the above research, which uses 90 training data consisting of 20 data tegal batik motifs, 50 data pekalongan batik motifs, and 20 cirebon batik motifs, as well as 40 test data, consisting of 10 data of original tegal batik motifs, 10 pekalongan batik motifs, can be concluded: 1) the accuracy of cirebon batik motifs and yogyakarta batik motifs using manhattan distance method is better than euclidean distance, which is 66%. 2) the value k=1 of manhattan distance and euclidean distance is the smallest is 54% for manhattan and 52% for euclidean distance. 3) predicted result of batik motif from pekalongan, which is 75% and followed by cirebon batik by 61.54%, then the prediction of batik tegal is 45.85% and then batik yogyakarta that is equal to 20%. for more research, researchers suggested that in testing it is recommended to use different object retrieval sizes and use different methods from the research created by current researchers in order to produce even better accuracy. references [1] f. a. etc, "pekalongan batik identification using the gray level co-occurrence matrix and probabilistic neutral network method," e-proceeding of engineering, vol. 6, p. 10234, 2019. 5 | vol. 4 no.1, january 2023 [2] h. c. d. a. s. a. a. halim, "image retrieval aplication using a combination of color moment and gabor texture methods," jsm stmik mikrisil, vol. 14, 2013. [3] a. k. a. a. susanto, image processing theory and application, yogyakarta: andi, 2012. [4] i. s. a. y. c. aj arriawati, "classification of texture images using k-nearest neighbor based on the characteristics extraction of the cookbook matrix method," diponegoro university, semarang. [5] n. s. a. a. w. b. arisandi, "introduction to batik motif using rotated wavelet filters and neural networks," juti, vol. 9, pp. 13-19, 2011. [6] s. d. cahyo, "comparative analysis of several edge detection methods using delphi 7," gunadarma university, depok, 2009. [7] c. c. a. t. b. o. c. regency, "casta and taruna, batik cirebon, world cultural heritage from indonesia, cirebon," cirebon, 2007. [8] d. p. pamungkas, "image extraction using glcm and knn methods to identify types of orchids (orchidaceae)," vol. 1, p. 2, 2019. [9] e. prasetyo, data mining processes data into information using matlab, yogyakarta: andi, 2014. [10] eliyani, the introduction of ripe papaya fruit levels using rgb color-based image processing with k-means clustering, lhokseumawe: lhokseumawe state polytechnic, 2013. [11] h. priyanto, digital image processing theory and real applications, bandung: informatics bandung, 2017. [12] h. wijayanto, "klasifikasi batik menggunakan metode k-nearest neighbour berdasarkan gray level co-occurance matrixes (glcm)," 2015. [13] j. ong, "implementation of the k-means clustering algorithm to determine president university's marketing strategy," scientific journal of industrial engineering, vol. 12, pp. 10-30, 2013. [14] i. m. johan wahyudi, "introduction to traditional fabric image patterns using glcm and knn," jtiulm, vol. 4, pp. 43-48, 2019. [15] m. s. etc, comparison of texture and color feature extraction for classification of lamongan batik, tuban, 2017. [16] a. h. a. a. p. h. rangkuti, "content-based drawing of batik," journal or computer science, vol. 10, pp. 925934, 2014. [17] a. p. d. &. w. r. triprasetyo, "application of trenggalek batik pattern recognition using sobel edge detection and kmeans algorithm," generation journal, vol. 2, pp. 25-32, 2018. [18] z. y. lamasigi, "dct for feature extraction based on glcm on batik identification using k-nn," jambura journal of electrical and electronics engineering, vol. 3, 2021. [19] n. l. w. s. r. ginantra, "detection of batik parang using the co-occurrence matrix and features geometric invariant moment with knn classification," lontar computer, vol. 7, p. 05, 2016. p-issn : 2715-2448 | e-issn : 2715-7199 vol.4 no.1 january 2023 buana information technology and computer sciences (bit and cs) 28 | vol.4 no.1, january 2023 detecting harmful activity in pilgrimage using deep learning musa dima genemo study program computing software engineering gumushane university, turkey email: musa.ju2002@gmail.com ‹β› abstract—cctv surveillance is the most extensively used intelligent latest innovation. the use of surveillance cameras has risen dramatically because of the con-venience of monitoring from anywhere and the reduction of crime rates in public areas. in this paper, we introduce the idea of bad vibe activity detection from live videos to enhance the security and safety of pilgrims. the proposed bad vibes activity recognition model is intended to be addressed in the most efficient manner possible using cutting-edge technologies such as tensorflow and keras. tensorflow was chosen because the project could be deployed to a mobile environment in the future with the possibility of extension of other areas such as airport security, bus stain, and public areas that may deserve special attention for security checks. we choose mediapipe ho-listic for employee bad vibe recognition in the model. keywords—artificial intelligence, classification, real-time object recognition, computer vision. abstrak—pengawasan cctv adalah inovasi cerdas terbaru yang paling banyak digunakan. penggunaan kamera pengawas telah meningkat secara dramatis karena kemudahan pemantauan dari mana saja dan pengurangan tingkat kejahatan di tempat umum. dalam makalah ini, kami memperkenalkan ide deteksi aktivitas getaran buruk dari video langsung untuk meningkatkan keamanan dan keselamatan jemaah. model pengenalan aktivitas getaran buruk yang diusulkan dimaksudkan untuk ditangani dengan cara seefisien mungkin menggunakan teknologi mutakhir seperti tensorflow dan keras. tensorflow dipilih karena proyek dapat diterapkan ke lingkungan seluler di masa mendatang dengan kemungkinan perluasan area lain seperti keamanan bandara, noda bus, dan area publik yang mungkin memerlukan perhatian khusus untuk pemeriksaan keamanan. kami memilih mediapipe ho-listic untuk pengenalan getaran buruk karyawan dalam model. kata kunci—artificial intelligence, klasifikasi, real-time object recognition, computer vision. i. introduction the use of surveillance cameras has risen dramatically because of the conven-ience of monitoring from anywhere and the reduction of crime rates in public areas. hajj is one of the five pillars of the islamic religion where the pilgrimage to the holy city of mecca in the kingdom of saudi arabia, which takes place in the last month of the year (hijri calendar) and which all muslims are obligated to make at least once throughout their lifetime if they can afford it [1]. before covid_19 emerged, 2.5 million people would travel every year to saudi arabia for hajj. due to this, the security of pilgrims needs special attention. new cut-ting-edge technology is required to ensure the safety of the people and the city where the hajj imitation takes place, as well as the detection of forbidden activi-ties and the carrying of prohibited things such as guns, flames, sharp metals, and the like. human activity recognition (har) is the ability to use sensors to ana-lyze human body indicators or motion and identify human actions or events [2]. har is regarded as a significant component in various scientific research set-tings, such as health [3], human-robot interaction [4], and security [5]. such technologies are in high demand during the hajj festival to safeguard pil-grims' safety. many people have become victims of the hajj scam in recent years, losing money, cellphones, and other valuables. nowadays terrorist acts pose the greatest danger to public safety [6]. prohibited things, such as carrying a gun, hurling a bomb, deceiving people, and threatening a suicide bombing, should be checked instantly. as a result, these challenges demand models that generate a warning or alarm. if accurate forecasts are provided in a timely manner, human lives can be saved by employing this newly introduced model. interleaved actions, such as throwing a stone at three walls (ramy al jamarat), which is also known as stoning the devil (sheytan) and running between mina and muzdalifah are a pillar of the hajj pilgrims. the stoning of the devil may cause prediction ambiguity, by throwing stones at people or away from the road and running between mina and muzdali-fah may cause prediction ambiguity with a sudden run. a recurrent neural net-work (rnn) is used to overcome activity overlapping difficulties. despite this, utilizing smart cctv surveillance reduces labor expenses while also increasing the security and safety of pilgrims. this study proposes a deep feature extraction mechanism for forbidden motion and activity identification 29 | vol.4 no.1, january 2023 to address the difficulties. we proposed a new model named l4-branched-action net. by using this new model, we extract features from the video frame and bels the activity to activity to their respective class like the need for special attention, or safe move. 64 layers of cnndeep architecture are used for feature extraction. to optimize the deep features that have been obtained, an aco feature selection technique is applied. by running convolution layers over pre-trained public data like the cifra-100. ii. method the proposed model will be presented in its entirety in this section. further-more, this section includes details of the proposed 64-layer classification algo-rithm. we used the cifar-100 dataset to train the proposed model, as well as feature extraction from the action recognition dataset using the proposed cnn architecture, feature selection using ant colony optimization (aco), and predic-tion using a variety of algorithms. for autonomous feature extraction from video frames and classifications events in the frame, a novel proposed 64-layer cnn architecture is used. the recommended l4-branchedactionnet's physical architecture is shown in fig.3 and fig.4. fig.3. structure of proposed model fig 4. video frame generation table 1. layer configuration of l4-branched action net lay er # layer name feature maps filter depth strid e 1 input 227 × 227 × 3 2 conv_1 55 × 55 × 96 11 × 11 × 3 × 96 [4 4] 3 relu_1 55 × 55 × 96 4 batch_norm_ 3 55 × 55 × 96 ….. fc_20 1 × 1 × 100 [1 1] same 62 prob 1 × 1 × 100 63 fc_21 1 × 1 × 100 [1,1] same 64 video description the data was collected using a script generated utilizing opencv and mediapipe holistic, as shown in fig.5 frames of data are recorded for each word caught. fig.4. key using open pose using mediapipe holistic point extraction [12] numpy array is used instead of pictures to hold video frames. we passed three major steps to train the model. the following are the details of the new model's operations steps. the first step the is conv layer; (1) in the conv layer the input x i−1 filter is computed using equation 1. where 𝑝𝑗 input channels and 𝑝^𝑗 represent the number of output channels. j represents several layers in the mode, fi filter. equation 2 is used to calculate the max pool in the pooling layer. where 𝑢, 𝑣 represents the matrix index of frame x𝑝, 𝑗-1and 𝑙, 𝑚 matrix index of the pooling window. it calculates the mean and variance in fragments. the mean is derived, and the features are separated using the standard deviation as follows. where 𝑤 is the number of feature maps in a batch. we used both relu and leaky_ relu in the proposed model. all numbers less than 0 are transformed to 0 by the standard relu, which is stated as [15]: for values less than zero, leaky relu has a small slope rather than zero. a leaky relu will have v = 0.01u when u is negative. cnn can further be learned in-depth from several works [16-19]. the second step is feature extraction;(2) for feature extraction from a video frame, the appropriate frame is retrieved. the proposed approach is intended to feature extraction from the deep-trained cnn pipeline. we trained the new model on public dataset t such as cifar100 [70] which contained images of 1000 classes. the trained network is then used for feature extraction on action recognition datasets and the fc_18 layer is chosen for features extraction. a total of 4096 features is attained per frame from the fc_18 layer. the prepared dataset contains a total of 13250 video 30 | vol.4 no.1, january 2023 frames. this makes the feature set dimension of all datasets 13250 × 4096. figure 5 illustrates the visualizations of the strongest feature maps at various convolution layers on l4branched-actionnet. fig.5. image visualizations of strongest feature maps at various convolution layers (a) conv_1, (b) conv_2, (c) conv_5, (d) g_conv_8, (e) conv_10. and the third step is (3) after interpreting the received result the extracted features are coded by applying entropy-coded aco optimization operation [25] using equation (5). where (x1-xn) represents the feature. we used aco for feature optimization based on the likelihood at a given point at a certain time. the last step is classification, in which aco-based chosen features are at the end passed to the predictor for categorization. several svm and knn versions are used to assess model performance. cub-svm emerges as the most effective as shown in table 2. table 2. performance of the model classifier sensitivity specificity precision measure percent lsvm 83.38 72.62 39.94 52.52 77.74 qsvm 89.11 91.53 61.79 76.01 86.14 fgsvm 57.29 51.78 25.02 32.80 54.39 mgsvm 90,52 92.35 62.56 76.58 86.28 cgcvm 68.47 64.45 31.80 41.75 66.33 csvm 96.33 95.59 76.61 88.08 92.99 for testing, we employed random selection using sklearn's train test function. following that, keras' callback functions were used to improve the training's efficiency. the accuracy of the test data is evaluated. we also used the public dataset on weizmann to compare our results to the current state of the art. the outcome is shown in table 3. table 3. performance evaluation on weizmann dataset method reference year accuracy dwt+knn [21] 2020 0.93 cnn+elm [22] 2020 0.94 gabor-ridgelet transform [23] 2020 0.93 lcf + msvm [22] 2021 0.95 ann [24] 2020 0.80 pcanet-xy-yt [25] 2021 0.91 ours (l4-branched-actionnet + entacs + cub-svm) 0.93 iii. results and discussion in the major goal of this study is to develop a cnn architecture that can recog-nize harmful actions during the hajj festival. then, the deep l4-branchedactionnet deep network proposed here is used to extract powerful features. the pretraining is carried out using a publicly available dataset, cifar-100. for testing, we employed random selection using sklearn's train test func-tion. following that, keras' callback functions were used to improve the training's efficiency to complete this design, many methods such as fine-tuning, add-ing and removing layers, and neurons were used. finally, the 64-layer architec-ture was proven it is the most efficient in terms of performance. tensor flow keras, opencv, and the numpy library were used in all the experiments in this. table 4. confusion matrix of csvm classifier sudden ran 0.92021 0.00 0.01 0.02 fighting 0.00 0.91221 0.00 0.01 throwing 0.01 0.01 0.90021 0.00 robbing 0.00 0.00 0.01 0.91002 sudden ran fighting throwing robbing iv. conclusion and recommendations detection of harmful vibes is critical for pilgrims' safety. to detect banned actions during the hajj festival, we utilized a 64-layer cnn network called l4-branched-actionnet. the model is evaluated on datasets that are freely availa-ble, such as the cifar-100 object detection dataset. the characteristics were retrieved and subsequently reduced using an entropy-coded aco. to evaluate model performance, several svm and knn versions are utilized. with an accu-racy of 0.91221, cub-svm emerges as the most effective. this work will be im-plemented on security personnel's mobile phones for convenient monitoring from any location in future work. references 1. r. k. tripathi, a. s. jalal, and s. c. agrawal, "suspicious human activity recognition: a review," artificial intelligence review, vol. 50, pp. 283-339, 2018 2. a. tapus, a. bandera, r. vazquez-martin, and l. v. calderita, "perceiving the person and their interactions with the others for social robotics–a review," pattern recognition letters, vol. 118, pp. 3-13, 2019. 3. a. ilidrissi and j. k. tan, "a deep unified framework for suspicious action recognition," artificial life and robotics, vol. 24, pp. 219-224, 2019. 4. konstantinidis, d., dimitropoulos, k., & daras, p. (2018). sıgn language recognıtıon based on hand and body skeletal data. 20183dtv-conference: the true vision apture, transmission and display of 3d video (3dtvcon). 5. s. j. elias, s. m. hatim, n. a. hassan, l. m. a. latif, r. b. ahmad, m. y. darus, and a. z. shahuddin, "face recognition attendance system using local binary pattern (lbp)," bulletin of electrical engineering and informatics, vol. 8, 2019. 6. a. krizhevsky, i. sutskever, and g. e. hinton, "imagenet classification with deep convolutional neural networks," 31 | vol.4 no.1, january 2023 advances in neural information processing systems, vol. 25, pp. 1097-1105, 2012 7. genemo, m. d. (2022). suspicious activity recognition for monitoring cheating in exams. proceedings of the indian national science academy, 1-10. 8. c. a. devine and e. d. chin, "integrity in nursing students: a concept analysis," nurse education today, vol. 60, pp. 133-138, 2018. 9. h. m. abdulghani, s. haque, y. a. almusalam, s. l. alanezi, y. a. alsulaiman, m. irshad, et al., "self-reported cheating among medical students: an alarming finding in a crosssectional study from saudi arabia," plos one, vol. 13, p. e0194963, 2018. 10. m. a. lewis and c. neighbors, "an examination of college student activities and attentiveness during a web-delivered personalized normative feedback intervention," psychology of addictive behaviors, vol. 29, p. 162, 2015. 11. rwth-phoenix-2014-t veri seti, https://wwwi6.informatik.rwth-aachen.de/~koller/rwthphoenix2014-t/ 12. s. balocco, m. gonzález, r. ñanculef, p. radeva, and g. thomas, "calcified plaque detection in ivus sequences: preliminary results using convolutional nets," in international workshop on artificial intelligence and pattern recognition, 2018, pp. 34-42 13. y. liu, x. wang, l. wang, and d. liu, "a modified leaky relu scheme (mlrs) for topology optimization with multiple materials," applied mathematics and computation, vol. 352, pp. 188-204, 2019 14. j. bouvrie, "notes on convolutional neural networks," neural nets, mit cbcl tech report, pp. 47-60, 2006. 15. y. li, z. hao, and h. lei, "survey of convolutional neural network," journal of computer applications, vol. 36, pp. 2508 2515, 2016. 16. a. divakaran, q. yu, a. tamrakar, h. s. sawhney, j. zhu, o. javed, et al., "real-time object detection, tracking and occlusion reasoning," ed: google patents, 2018. 17. a. booranawong, n. jindapetch, and h. saito, "a system for detection and tracking of human movements using rssi signals," ieee sensors journal, vol. 18, pp. 2531-2544, 2018. 18. a. b. mabrouk and e. zagrouba, "abnormal behavior recognition for intelligent video surveillance systems: a review," expert systems with applications, vol. 91, pp. 480491, 2018. 19. krizhevsky and g. hinton, "learning multiple layers of features from tiny images (technical report)," university of toronto, 2009 20. d. k. vishwakarma, "a two-fold transformation model for human action recognition using decisive pose," cognitive systems research, vol. 61, pp. 1-13, 2020. 21. m. a. khan, y.-d. zhang, s. a. khan, m. attique, a. rehman, and s. seo, "a resource conscious human action recognition framework using 26-layered deep convolutional neural network," multimedia tools and applications, pp. 1-23, 2020. https://wwwi6.informatik.rwth-aachen.de/~koller/rwth-phoenixhttps://wwwi6.informatik.rwth-aachen.de/~koller/rwth-phoenix p-issn : 2715-2448 | e-isssn : 2715-7199 vol.3 no.1 january 2022 buana information tchnology and computer sciences (bit and cs) 17 | vol.3 no.1, january 2022 web-based public complaints information system for subang city emy lenora tatuhey1 study program, information system ,stimik sepuluh nopember jayapura emytatuhey@gmail.com tukino 2 study program, information system faculty of computer science, universitas buana perjuangan karawang, indonesia tukino@ubpkarawang.ac.id ‹β› irpan hilmi 3 study program, information system faculty of computer science, universitas buana perjuangan karawang, indonesia irpanhilmi@ubpkarawang.ac.id abstra —keluhan masyarakat pada saat ini terbatas oleh media sebagian masyarakat merasa sulitnya untuk menyampaikan atau melaporkan peristiwa yang terjadi diwilayah tersebut. pengaduan keluhanpun masih sepenuhnya belum menggunakan media teknologi informasi hal ini membuat masyarakat sulit untuk memberikan saran atau menyampaikan keluhannya. dari permasalahan tersebut penulis mengembangkan sistem informasi untuk proses pengaduan masyarakat dengan metode waterfall. tujuan dibuatkannya sistem ini yaitu untuk membuat website desa jatibaru, membuat website pengaduan bagi masyarakat, sehingga keluhan yang ada dimasyarakat bisa tersampaikan dengan baik melalui website pengaduan, masyarakat dapat melihat kembali pengaduan untuk mengetahui apakah pengaduan sudah diproses atau belum. pengembangan sistem menggunakan bahasa pemograman php dan html. hasil akhir dari penelitian ini menghasilkan sebuah sistem informasi pengaduan masyarakat supaya aspirasi dari masyarakat tersampaikan dengan baik untuk pemerintah desa, semua pengaduan akan dibahas oleh aparatur desa untuk meminimalisir pengaduan dari masyarakat dan dengan adanya sistem ini juga dapat membantu memonitoring pengaduan yang disampaikan masyarakat. kata kunci : pengaduan masyarakat, keluhan, sistem informasi abstract—the media currently limit public complaints, and some people find it difficult to convey or report events in the area. complaints are still not fully using information technology media. this makes it difficult for the public to provide suggestions or submit complaints. the authors developed an information system for the public complaint process using the waterfall method from these problems. this system aims to create a jatibaru village website and create a complaint website for the community so that complaints in the community can be conveyed properly through the complaints website. the public can review the complaints to determine whether the complaints have been processed or not. system development using the php and html programming languages. the final result of this study resulted in a public complaint information system so that the aspirations of the community were properly conveyed to the village government, all complaints would be discussed by the village apparatus to minimize complaints from the community, and with this system, it can also help monitor complaints submitted by the community. keywords: public complaints, complaints, information systems i. introduction a government is expected to have service media that can be accessed via the internet to accommodate input from the community. the author proposes to make a web-based jatibaru village community complaint service. make it easier for the community to give or complain suggestions and criticisms about the problems in jatibaru village because these problems can become obstacles to the progress of a village. all residents of jatibaru village can access the complaint system to provide suggestions and criticisms for the village government to be more advanced and better in the future; aspirations in the form of suggestions and criticisms are a form of attention from the community to the village government. if the community wants to convey their aspirations, they can open the jatibaru village website and enter the username and password that the community already has. after entering the system, there will be several menus, one of which is a menu for complaints or aspirations from the community to the village. the community needs to fill out the complaint form. and can submit more than one complaint. the submitted complaint form will be stored and recapitulated by the village. every week the village holds weekly meetings, which are held on wednesdays, to discuss complaints submitted by the community to minimize negative things that arise in the community. the village head also evaluates the rt or kadus who are less active in attending weekly meetings held at the jatibaru village hall. 18 | vol.3 no.1, january 2022 ii. method a. data collection techniques in this process, in-depth research is carried out on the data needed during the process of making the system to be designed. following are the stages of data collection. a) literature study looking for theoretical references related to the research topic raised, theories obtained from several literature sources. there are several sources of literature that are often used in the form of books, journals, theses. b) observation the purpose of the observation is to observe the research site in order to get the information needed by the researcher, by going directly to the object being studied, namely the jatibaru village office, ciasem subang. c) interview in this interview, conducting a dialogue by asking questions to the village apparatus and the jatibaru village community aims to find out the process of submitting complaints so that it becomes material for building a new system. b. system development method the waterfall method is a method commonly referred to as the classic life cycle, describing a systematic approach to sequential software development. [9] the following are the stages of the waterfall model: figure 1 iterative waterfall[12] 1. system information and engineering modeling this stage starts from gathering requirements to be applied to the device system that will be developed. 2. software requirements analysis this stage is to collect the required needs with incentives to be understood by users, these needs are intended for users as system users later. 3. design this stage does the design of the interface design requirements for the system to be developed, the interface design needs to be done when you want to do system development. 4. coding the stages of making this program must be implemented using software. 5. testing this stage of system testing is carried out logically and functionally which is used to determine the parts of the system that have been previously tested. this testing is to minimize errors that occur in the program. 6. maintenance the maintenance stage or system maintenance needs to be done because the system needs can continue to be improved after use. iii. results and discussion in this study, the collection of survey requirements aims to analyze existing problems in village government and evaluate the standard application process. in this case, it is a systematic analysis. present and find solutions to village and community leaders' problems and needs to find out user desires for applications used with the waterfall method as system development. here are the steps: a. current system analysis at the jatibaru village agency currently, it is still done manually or by greeting with the village apparatus, not a few people want to express opinions and input for the jatibaru village government to be better, some problems or obstacles that occur, namely regarding the process of submitting complaints, they still often experience problems. impasse, such as the village head and village apparatus, which are often difficult to find, the community is reluctant to have direct dialogue when conveying their aspirations to the village apparatus, and also both village officials and the community are often constrained by time when they want to go to the village office in carrying out discussions, both residents who want to express their aspirations as well as village officials who want to respond to all community complaints. figure 2 running system 19 | vol.3 no.1, january 2022 b. system design after knowing the description that has been running and the analysis results to determine user satisfaction regarding the system developed at the jatibaru village office, ciasem subang next is the development and design of the system. the following are the stages in making the system: 1. system proposed the proposed system aims to identify and evaluate problems found during research. the proposed system consists of analyzing problems and analyzing needs. therefore, the author gives the jatibaru village agency an idea to handle the above problems. below is the proposed system procedure for jatibaru village, ciasem subang. after analyzing the current system in jatibaru village, a system is proposed that will provide online complaint services. figure 3 system flowchart in the flow above, it can be seen that the community will start the complaint service by opening the website that has been built. then the community can enter their username and password to be able to log in to the home page then choose the complaint form and fill out the complaint, after that the admin receives the complaint and the village head discusses to provide a reply to the complaint submitted by the community. the village head confirms to the admin to reply to the complaint, then the officer in charge of the complaint can upload the progress of the settlement that is being carried out, after the complaint is processed the community can view the complaint again [10]. 2. use case diagram use case diagrams are diagrams that describe the scenario of the system that will be created and explain between the actors and the activities that will be carried out on the system that has been built. the use case diagram below describes the overall activities carried out by each user. figure 4 usecase diagram 3. activity diagram this activity diagram illustrates what the public does in making complaints. by logging in first, then if the data validation is correct, the system will display the user's main page. next, the user will select the complaint menu, the system displays the complaint form, then the user fills out the complaint if the complaint form has been filled in, the system will return to the main page and the complaint data is stored in the database. users can also add complaints if they want to make more than one complaint. 4. sequence diagram in this complaint input sequence diagram, it describes the lifeline between objects when the user inputs a complaint. below is the sequence diagram for the complaint input: figure 6. sequence diagram of complaint input 20 | vol.3 no.1, january 2022 5. class diagram class diagram is a diagram that describes the form of the system in the process of class definition on the system to be built. figure 7. class diagram c. system implementation in the implementation of this interface, it displays a system display that has been created and run using google chrome and xampp as the web server. here is the interface implementation view: figure 8. login page figure 9. registration page figure 10. forgot password page figure 11. complaint input page figure 12. complaint review page figure 13. complaint responding page d. system evaluation this web-based public complaint information system aims to create a website where people can communicate their wishes to village officials, the wishes that exist in the community are solely for the progress of the village itself both in the field of development, facilities and infrastructure 21 | vol.3 no.1, january 2022 in jatibaru village, every user who participating are granted access with privileges granted by the system, therefore every connection with the system already has a limit for each user. the results of testing the system using the black box testing method to determine its functionality so that it can be concluded that the resulting system functions as expected. the system that was built also helps the village head and village parties to know what is being communicated by the community. iv. conclusion based on the research results in making the final project on a web-based public complaint information system at the jatibaru ciasem subang village office, the researchers got several conclusions, namely as follows: 1. with the creation of a web-based public complaint system, the complaint process can be carried out easily and relevantly for the jatibaru village community. this system also has a complaint info feature to make it easier for the public to see the progress of the complaints submitted. 2. complaints submitted will be discussed by village officials and village heads so that they can be used as evaluation material for village officials to be even better. reference [1] d. haryanto and a. nasihin, “sistem informasi kearsipan surat masuk surat keluar di stikes mitra kencana kota tasikmalaya,” j. tek. inform., vol. 6, no. 2, pp. 22–30, 2018, [online]. available: http://jurnal.stmik-dci.ac.id/index.php/jutekin/. [2] b. huda, “sistem informasi data penduduk berbasis android dan web monitoring studi kasus pemerintah kota karawang (penelitian dilakukan di kab. karawang),” buana ilmu, vol. 3, no. 1, pp. 62–69, 2018, doi: 10.36805/bi.v3i1.456. [3] s. w. mursalim, “analisis manajemen pengaduan sistem layanan sistem aspirasi pengaduan online rakyat (lapor) di kota bandung,” j. ilmu adm. media pengemb. ilmu dan prakt. adm., vol. 15, no. 1, pp. 1–17, 2018, doi: 10.31113/jia.v15i1.128. [4] a. s. cahyono, “pengaruh media sosial terhadap perubahan sosial masyarakat di indonesia,” j. ilmu sos. ilmu polit. diterbitkan oleh fak. ilmu sos. polit. univ. tulungagung, vol. 9, no. 1, pp. 140–157, 2016, [online].available:http://www.jurnalunita.org/index.p hp/publiciana/article/download/79/73. [5] e. widyawati, “rancang bangun aplikasi kependudukan berbasis web di desa kedungrejo waru-sidoarjo,” j. manaj. inform., vol. 6, no. 1, 2016. [6] s. anton, i. c. alex, and n. kristiawan;, “sistem penilaian kinerja pegawai dalam pelayanan nasabah pada … (sujarwo dkk.),” penilaian, sist. pegawai, kinerja nasabah, pelayanan kinerja, abstr. lang. unified model., pp. 270–275, 2019, [online]. available: http://undhari.ac.id/jurnal/index.php/simtika/article/vie w/15. [7] c. surya and s. sara, “jaringan sistem informasi robotik vol. 2, no. 02, september 2018,” jar. sist. inf. robot., vol. 2, no. 02, pp. 115–129, 2018. [8] d. utama, a. johar, and f. f. coastera, “minuman restaurant berbasis client server dengan p latform android,” pp. 288–300, 2016. [9] a. galih pradana and s. nita, “rancang bangun game edukasi ‘ amudra ’ alat musik daerah berbasis android afista galih pradana sekreningsih nita,” semin. nas. teknol. inf. dan komun., vol. 2, no. 1, pp. 77–80, 2019. [10] d. kurniawan, a. l. hananto, and b. priyatna, “modification application of key metrics 13x13 cryptographic algorithm playfair cipher and combination with linear feedback shift register (lfsr) on data security based on mobile android,” int. j. comput. tech.-–, vol. 5, no. 1, pp. 65–70, 2018. p-issn : 2715-2448 | e-isssn : 2715-7199 vol.3 no.2 july 2022 buana information tchnology and computer sciences (bit and cs) 35 | vol.3 no.2, july 2022 signature verification using the k-nearest neighbor (knn) algorithm and using the harris corner detector feature extraction method aang alim murtopo 1 study program technical information stmik ymi tegal aang.alim@gmail.com bayu priyatna 2 study program information system universitas buana perjuangan karawang bayu.priyatna@ubpkarawang.ac.id ‹β› rini mayasari 3 study program information system universitas singaperbangsa karawang rini.mayasari@staff.unsika.ac.id abstrak— keamanan proses transaksi sangat penting di zaman sekarang ini. tanda tangan dapat digunakan sebagai sarana penjamin keamanan suatu transaksi selain sidik jari. namun, ancaman pemalsuan tanda tangan bagi mereka yang menggunakan tanda tangan sebagai keamanan masih sangat tinggi dan sering terjadi. dalam penelitian ini, kami akan memverifikasi keaslian tanda tangan dan mengujinya menggunakan algoritma knearest neighbor (knn) dan metode ekstraksi fitur harris corner. ada dua macam perhitungan jarak yang akan digunakan pada algoritma k-nn yaitu dengan menghitung jarak dari euclidean distance dan manhattan distance. kata kunci— tanda tangan, verifikasi, knn (k-nearest neighbour), harris corner detector, euclidean distance, manhattan distance. abstract — the security of the transaction process is very important in this day and age. signatures can be used as a means of guaranteeing the security of a transaction other than fingerprints. however, the threat of signature forgery for those who use signatures as security is still very high and frequent. in this research, we will verify the authenticity of a signature and test it using the k-nearest neighbour (knn) algorithm and the harris corner feature extraction method. there are two kinds of distance calculations that will be used in the k-nn algorithm, namely by calculating the distance from euclidean distance and manhattan distance. the k value at knn taken is at k = 1, k = 3, and k = 5. keywords— signature, verification, knn (k-nearest neighbour), harris corner detector, euclidean distance, manhattan distance. i. introduction security in transactions needs to be given important attention today. there are several ways that can be done for transaction security measures such as fingerprints and pins are examples of frequently used security. one of the well-known safeguards is to use a signature. signatures are considered easier to use, cheaper, and quite effective. however, signature forgery is still common and poses a security threat to signature users. signature verification is a way to find out if the signature is genuine or fake. signature verification is divided into two forms, namely on-line and off-line signature verification. on-line means, when the signature is taken, recording of the time, pressure, and others. in off-line verification, the signature is only taken from a photo or image of that signature. the problem that often occurs in signature verification is the difference between the tools used to write signatures and the image retrieval process, image resolution, images that contain noise, and others. this signature itself is a handwriting that has a unique character and each person must have a difference. we often encounter unreadable signatures, however, signatures can be read as images and recognized by computers [1]. because of its uniqueness, the signature can be used as a security system and identifier of a person's identity. various studies that have been done before, have used various methods and feature extraction in this signature verification. among them, research [2] was conducted to verify signatures online using the feed forward back propagation error neural network and the discrete wavelet transform feature extraction method. research [3] used the euclidean distance model and geometric centre method for feature extraction. research [4] with k-nearest neighbour and gabor wavelet as characteristic extraction. research [5] used pca classification and multilayer feed forward artificial neural network and for its extraction using fourier descriptor and chain codes. research [6] used the euclidean distance classification and a new extraction method that is dividing into several images. 36 | vol.3 no.2, july 2022 from some of the studies above, these researchers have attempted to find a method of classification and feature extraction that can produce the highest level of accuracy. this level of accuracy is very important to avoid signature forgery so as to provide a sense of security in transactions and everything related to the economic sector. this research will use the k-nearest neighbour (knn) algorithm and the harris corner method for its feature extraction. from several previous studies, knn is considered capable of classifying well for difficult images, such as in the study [7] "automatic medical image classification and abnormality using k-nearest neighbour", knn classifies medical images with an accuracy of 80% and greater when compared to svm linear and rbf kernel. harris corner can be used for grayscale images and produces a more consistent extraction value from distorted images. as in research [8] "harris operator corner detection using sliding window method", with the harris corner method palms can be detected with an accuracy of 97.5%. ii. method 2.1 . study of literature research [2], which has been done, namely online signature verification, the classification algorithm used is feed forward back propagation error neural network and feature extraction of discrete wavelet transform produces an accuracy of 95%. using a sample of 100 signatures consisting of 10 original signatures and 10 fake signatures for each person. this sample was taken from 5 people who gave signatures. research [3] with the euclidean distance model classification algorithm and geometric center feature extraction on signature verification, resulted in a random far of 2.08%, simple 9.75% and 16.36% skilled forgeries. meanwhile, the frr was 14.58%. signatures tested were as many as 21 original signatures and 30 fake signatures. from these signatures, 9 original signatures were found. research [4] carried out from signature verification, using the nearest neighbour classification algorithm and using gabor wavelet feature extraction. this study resulted in verification with accuracy ≥ human accuracy in carrying out signature verification with the smallest far and frr values in this study were 22.5% and 15.5%. research [5] identified signatures based on fourier descriptor and chain codes and produced far = 2.6% and frr = 1.6% in the verification process. this study uses the pca algorithm and ann feed forward for classification. meanwhile, for feature extraction using fourier descriptor and chain codes. research [6] the approach used for feature extraction is to divide the image into rectangles based on the midpoint of gravity of the signature. the classification uses euclidean distance. the result is 0% random far, 0% simple and 1% skilled. meanwhile, the frr is 0.5%. a) k-nearest neighbour (knn) k-nearest neighbour is a non-parametric classification even though it is simple but it is one of the popular and well-known and often used. the key to this method is the user-defined k parameter. when k is selected and given the x pattern, assign the pattern to the class that has the greatest number of knearest neighbour (a calculation to measure the distance to the nearest neighbour), because k-nn uses a calculation by determining the closest distance. the calculation of the distance used in the k-nn [9]: i. euclidean distance ii. manhattan distance b) harris corner harris corner detector (harris angle detector) is a point (angle) detector that is often used because it is able to produce consistent values despite rotation, scale, lighting variations and noise. harris angle detector based on the autocorrelation function of the local signal which calculates the local change of the signal. this detector also functions to detect local gradients in the horizontal and vertical directions at each surrounding point, the aim is to find the image value whose intensity varies from the two directions. harris corner based on: 𝑆𝑖𝑗 = ∑ ∑ 𝑊𝑚𝑛 [ ℎ𝑚𝑛 2 ℎ𝑚𝑛𝑣𝑚𝑛 ℎ𝑚𝑛𝑣𝑚𝑛 𝑣𝑚𝑛 2 ] 𝑗+𝐷 𝑚=𝑗−𝑑 𝑖+𝐷 𝑚=𝑖−𝐷 where is 𝑆𝑖𝑗 calculated in the area of measure (2d + 1) x (2d + 1) around position (i, j). ℎ𝑚𝑛 represents the derived filter response horizontally, 𝑊𝑚𝑛 on vertical, and 𝑊𝑚𝑛 is the weight that reduces the impact of the position. the following is the harris corner detection algorithm [10]: 1. calculate the x and y derivatives of the figure. 37 | vol.3 no.2, july 2022 2. calculate the derivative of each pixel 3. calculate the product of the derivative of each pixel. 4. matrix form. 5. calculate the detection response in each pixel. 6. threshold response value. c) recall, precision, true negative rate and accuracy recall to find out the answers to the system obtained from: precision to determine the accuracy of the system to recognize the authenticity of signatures obtained from: accuracy to determine the performance of the feature extraction and classification used. table 1. confusion matrix total population true condition positive condition negative condition prediction condition predicted condition positive true positive false positive predicted condition negative false negative true negative true positive (tp) is when the program recognizes the authenticity of a signature image, it indicates that the signature image is genuine (true). false positives (fp) are when the program mistakenly recognizes the fake signature image as the original signature image. false negative (fn) is when the program mistakenly recognizes the original signature image as a fake signature image. true negative (tn) is when the program recognizes a fake signature image as a fake (true) signature image. 2.2 data collection signature image data, obtained from the scanning process and measuring 400x400 pixels totalling 300 images consisting of 150 original signatures and 150 fake signatures. these signatures were obtained from 10 different people. 2.3 design this research step is depicted in figure 1. verification of the signature starts from inputting the image of the signature to be tested, in the form of a grayscale image measuring 400x400 pixels, then extracting it to get the characteristics of the image. the results of this feature extraction in the form of coordinates that show the location of the angle are then stored with the knn model of the grayscale image feature extraction to be an example. the extraction results are then calculated the distance with the extracted samples one by one. after the distance calculation results are known, they are sorted from smallest to largest. the signature to be verified by category/label (fake or real), selected from a predetermined radius, will be tested. figure 1. design diagram iii. results and discussion this signature verification system, its appearance can be seen in figure 2 below. this signature verification system will be carried out in two ways for calculating the distance, namely euclidean distance and manhattan distance. in figure 2, to verify the signature by pressing the test image button to input the image to be tested. then press the original image button or the fake image button to input the image that will be used as training data. original and fake training data signature image is required for verification. after that the training data image is entered, press the verification button, the results will appear in the massage box. 38 | vol.3 no.2, july 2022 figure 2. display of the verification system test result the test carried out aims to determine the results of the level of accuracy. in figure 3 below, an example of a signature that will be used in research: figure 3. sample signature in figure 4 below, are the results of the signature test image. the signature image that will be used is obtained from the results of the scan process without resizing, totalling 300 image data. the data is divided into 200 images of training data and 100 images as test data. the signature image which is the training data will not be used in the test data and vice versa. when testing, training data and test data will go through a feature extraction process using the harris corner method. the results of the coordinates generated in the feature extraction process will be calculated the distance to knearest neighbour. testing on the same signature image model using three kinds of variable k at knearest neighbour, namely k = 1, k = 3, k = 5 for the calculation of euclidean distance. figure 4. image testing results table 2. signature testing accuracy k = 1 euclidean distance model type precision (%) recall (%) true negative rate (%) accuracy (%) 0 0 0 100 50 1 100 80 100 90 2 50 100 0 50 3 50 100 0 50 4 50 100 0 50 5 50 100 0 50 6 50 100 0 50 7 0 0 100 50 8 0 0 100 50 9 50 100 0 50 testing with k = 1 in table 2 using euclidean distance calculations produces an average precision of 40%, an average recall of 68%, an average true negative rate of 40%, and an average accuracy of 54%. table 3. signature testing accuracy k = 3 euclidean distance model type precision (%) recall (%) true negative rate (%) accuracy (%) 0 0 0 100 50 1 50 40 60 50 2 100 20 100 60 3 50 100 0 50 4 50 100 0 50 5 50 100 0 50 6 50 100 0 50 7 0 0 100 50 8 0 0 100 50 9 50 100 0 50 testing with k = 3 in table 3 using euclidean distance calculations produces an average precision of 40%, an average recall of 56%, an average true negative rate of 46%, and an average accuracy of 51%. table 4. signature testing accuracy k = 5 euclidean distance model type precision (%) recall (%) true negative rate (%) accuracy (%) 0 0 0 100 50 1 66.66 40 80 60 2 0 0 100 50 3 50 100 0 50 4 50 100 0 50 5 50 100 0 50 6 50 100 0 50 7 0 0 100 50 8 0 0 100 50 9 50 100 0 50 39 | vol.3 no.2, july 2022 testing with k = 5 using the euclidean distance calculation in table 4 produces an average precision of 31.66%, an average recall of 54%, an average true negative rate of 48%, and an average accuracy of 51%. table 5. signature testing accuracy k = 1 manhattan distance model type precision (%) recall (%) true negative rate (%) accuracy (%) 0 0 0 100 50 1 0 0 100 50 2 50 100 0 50 3 50 100 0 50 4 50 100 0 50 5 50 100 0 50 6 50 100 0 50 7 0 0 100 50 8 0 0 100 50 9 50 100 0 50 testing with k = 1 using the manhattan distance calculation in table 5 produces an average precision of 30%, an average recall of 60%, an average true negative rate of 40%, and an average accuracy of 50%. table 6. signature testing method average precision (%) recall average (%) average true negative (%) average accuracy (%) euclidean k = 1 40 68 40 54 euclidean k = 3 40 56 46 51 euclidean k = 5 31.66 54 48 51 manhattan k = 1 30 60 40 50 from the results of tests carried out in table 6, it is known that the k = 1 value of the euclidean distance calculation is the best with an accuracy of 54%. however, the value of k = 1 is prone to noise. with a value of k = 1, if the distance to the noise image is the smallest, then the image that is verified will be considered the same type of image as the noise image. this is indicated by the smallest true negative rate with a value of 40% (equal to k = 1 manhattan distance calculation). a small true negative rate indicates the system has probably made a mistake by accepting a fake signature as the original signature is large enough. the value of true negative rate depends on the similarity of the fake signature and image noise in the sample image. the greater the similarity of fake signature and image noise in the sample image, from the above test, it can be seen that the euclidean distance calculation is better than manhattan distance. this can be seen from the greater precision, recall, and accuracy euclidean distance values compared to manhattan distance. iv. conclusion signature verification can be done to determine the authenticity of the signature. signature verification is done based on the angle found and then applies the knearest neighbour algorithm. the small k value in the k-nearest neighbour algorithm has a tendency to accept the fake signature image as the original signature image is greater than the larger k value. euclidean distance calculation is better used for signature verification than manhattan distance calculation. at k = 1, the accuracy of the euclidean distance is 54% while the manhattan distance is 50%. v. suggestion as for suggestions that are useful for further research, namely: 1. further research can take a signature image using a pen tablet to reduce noise in the signature image. 2. future research can reproduce the signature image that will be used as an example. 3. signature verification can be done using other algorithms. a. references b. [1] c. oz, "signature recognition and verification with artificial neural network using moment invariant method," in international symposium on neural networks, china, 2005. [2] m. m. fahmy, "online handwritten signature verification system based on dwt features extraction and neural network classification," ain shams engineering journal, vol. 1, p. 59– 70, 2010. [3] y. s. r. d. p. b. banshider majhi, "novel features for off-line signature verification," international journal of computers, communications & control (ijccc), vol. 1, pp. 17-24, 2006. [4] m. r. p. h. r. p. mohamad-hoseyn sigari, "offline handwritten signature identification and verification using multi-resolution gabor wavelet," ciit international journal of biometrics and bioinformatics, vol. 5, pp. 234248, 2011. [5] m. a. r. t. e. d. a. s. ismail a. ismail, "an efficient offline signature identification method based on fourier descriptor and chain codes," international journal of biomedical engineering and technology, vol. 5, pp. 1-10, 2011. 40 | vol.3 no.2, july 2022 [6] p. i. s. s.a. daramola, "novel feature extraction technique for off-line signature verification system," international journal of engineering science and technology, vol. 2, pp. 3137-3143, 2010. [7] m. k. rakesh ramteke, "automatic medical image classification and abnormality detection using k-nearest neighbour," international journal of advanced computer research, vol. 2, pp. 190-196, 2012. [8] r. d. g. s. jyoti malik, "harris operator corner detection using sliding windows method," international journal of computer applications, vol. 22, 2011. [9] a. kadir, image processing theory and application, yogyakarta: andi offset, 2013. [10] s. j. prince, computer vision: models, learning and inference, cambridge: cambridge university press, 2012. 47 | vol.3 no.2, july 2022 p-issn : 2715-2448 | e-isssn : 2715-7199 vol.3 no.2 july 2022 buana information tchnology and computer sciences (bit and cs) automatic face mask detection on gates to combat the spread of covid-19 musa dima genemo study program computing (software engineering) gumushane university musa.ju2002@gmail.com ‹β› abstract— the covid-19 pandemic has spread across the globe, hitting almost every country. to stop the spread of the covid-19 pandemic, this article introduces face mask detection on a gate to assure the safety of instructors and students in both class and public places. this work aims to distinguish between faces with masks and without masks. a deep learning algorithm you only look once (yolo) v5 is used for face mask detection and classification. this algorithm detects the faces with and without masks using the video frames from the surveillance camera. the model trained on over 800 video frames. the sequence of a video frame for face mask detection is fed to the model for feature acquisition. then the model classifies the frames as faces with a mask and without a mask. we used loss functions like generalize intersection of union for abjectness and classification accuracy. the datasets used to train the model are divided as 80% and 20% for training and testing, respectively. the model has provided a promising result. the result found shows accuracy and precision of 95% and 96%, respectively. results show that the model performance is a good classifier. the successful findings indicate the suggested work's soundness. keywords: accuracy, computer vision, classification, face recognition, surveillance 1. introduction the covid-19 pandemic has spread across the globe, hitting almost every country. the epidemic was initially discovered in wuhan, china, in december 2019. the codi-19 epidemic has allowed us to pave the road for digital learning to be introduced. to stop the spread of the covid-19 pandemic, this article introduces face mask detection at university gates. face recognition applications employ various deep learning algorithms to look for human faces within larger pictures that often include non-facial items such as landscapes, buildings, and other human body parts such as feet or hands. the search for human eyes, which is one of the simplest characteristics to identify, is often where face identification algorithms get started. after that, the algorithm could make an effort to recognize the iris, the mouth, the nose, the nostrils, and the eyebrows. after the algorithm has concluded that it has located a facial area, it next conducts further tests to ensure that it has identified a face in the image [1]. the algorithms need to be trained on big data sets that include hundreds of thousands of examples of both face masks and without face mask pictures. this will assist guarantee that the results are accurate. the training helps the algorithms get better at determining whether a picture contains a face mask. where those faces are located within the image. knowledge based, feature-based, template-matching or appearance based approaches are some of the strategies that may be used in face identification [2]. the deep learning algorithm yolo is introduced that detects the objects efficiently. the authors do research and improve its version to v5 as shown in figure 1. each one has positive and negative aspects to consider. the problem occurs that we need an algorithm that accurately recognized the faces. in this paper, we have highlighted the issue that the algorithm accurately recognized the faces with and without masks. to build a model that detects face masks we used a new deep learning algorithm yolo v5. this algorithm efficiently resolves the problem and accurately recognizes the faces. figure 1: yolo detection process [4] the following is the order in which the manuscript is written. the introductory section and the literature review overview section are included in sections 1 and 2, respectively. the proposed approach is covered in section 3. section 4 shows the findings and explanation of the performance evaluation 48 | vol.3 no.2, july 2022 experiments. the paper's conclusions are presented in section 5. 2. literature review knowledge-based approaches, also known as rule-based methods, characterize a face by adhering to certain guidelines. the difficulty of formulating rules that are clearly stated is one of the drawbacks of using this method. noise and light may have a detrimental effect on facial recognition techniques known as feature invariant methods [3]. these techniques employ distinguishing characteristics of a person's face, such as their eyes or nose, to identify a person's face. the detection of faces using template-matching approaches involves comparing pictures with conventional face patterns or traits that have been recorded in the past and connecting the two to establish a relationship between the three. the problem with these approaches is that they do not account for differences in position, size, or form. finding the important attributes of face photos requires statistical analysis and machine learning, both of which are used by appearance based approaches [4]. this approach, which is also used in the process of feature extraction for face recognition, is broken down into many sub-methods. the following is a list of some of the more specialized methods that are used in face detection: taking off the backdrop image. for instance, if a picture has a backdrop that is a single color, or if it has a background that is pre-defined and unchanging, then removing the background from the image may assist show the facial borders. sometimes the color of the subject's skin may help identify faces in color photographs; however, this is not guaranteed to work with every complexion. the use of motion to identify faces is yet another possibility. because a face is generally always moving during the real time video, users of this approach are required to determine the region of the face that is moving. this approach has a few drawbacks, one of which is the potential for misunderstanding with other moving objects in the background. a complete technique for detecting faces may be created by combining some of the tactics described in the previous paragraph [5]. face recognition in photographs can be challenging because of the many variables that can affect the process, differences in camera gain, lighting conditions, and image quality, as well as differences in attitude, emotion, location, orientation, skin color, and pixel values, the existence of spectacles or facial hair, and the photograph's orientation are all factors to consider. deep learning, which has been used to make advancements in face identification in recent years, has the benefit of greatly outperforming standard computer vision approaches. these advancements have been made in recent years. artificial intelligence detects objects. artificial intelligence required a large amount of data. then as a solution authors use yolo, however, a single yolo version did not detect a large number of objects accurately [6]. in [7] authors, highlight the issue of car number plate detection. authors use the yolo algorithm that does not accurately detect all number plates. the authors in [8], and [9] highlight the issue of face recognition. many other authors use other ml techniques; however, that techniques fail to capture the facial expressions. the accuracy of the algorithms is not good. the authors resolved the problem of detecting apple plants’ pictures. many authors use deep learning algorithms that fail to recognize that yolo v4 is used to detect pictures whose accuracy is also low [10-11]. paul viola and michael jones, computer vision researchers, developed a system in 2001 that could accurately recognize faces in real-time. these developments brought about significant advancements in the face detection approach [1, 12]. the viola-jones framework relies on the concept of teaching a model to recognize what constitutes a face and what does not constitute a face. once the model has been trained, it will extract certain features, which will then be saved in a file. this will allow the features extracted from fresh photos to be compared with the features that were previously recorded at different phases [13-15]. if the picture being analyzed is successful in passing through each step of the comparison of its features, then a face has been identified and the processes may continue. even though it is still widely used, the viola-jones framework has certain shortcomings when it comes to the recognition of faces in real-time applications. for instance, the framework may not function correctly if a face is obscured by something like a mask or a scarf. likewise, if a face is not orientated correctly, the algorithm might not be able to locate it [16]. other algorithms have been created to help improve processes and eliminate the disadvantages of the viola-jones framework, such as the region-based convolutional neural network (r-cnn) and single shot detector (ssd). this has helped to improve the overall quality of the process [3]. in the field of image identification and processing, an artificial neural network known as a convolutional neural network (cnn) is a form of neural network that was developed for the sole purpose of handling pixel input. to localize and categorize the objects seen in pictures, r-cnn will provide region recommendations based on a cnn framework. ssd requires only one shot to recognize several objects within an image, compared to two shots required by region proposal network-based techniques like r-cnn [27,28]. the first shot is used to generate region proposals, and the second shot is used to detect the object associated with each proposal [2]. as a result, ssd is much quicker than r-cnn. the covid-19 outbreak has quickly wreaked havoc on our day-to-day lives, impeding the flow of commerce and travel across international borders. protecting one's face by using a face mask has emerged as the new standard practice. in [17, 18] the not too distant future, many suppliers of public services will need their customers to wear masks that are suited for the environment to get their services. as a result, the identification of face masks has developed into an essential obligation to support international culture. in [19], [20] using several essential machine learning technologies such as tensorflow, keras, opencv, and scikit-learn, the method presented in this research offers a straightforward approach to accomplishing this goal. the 49 | vol.3 no.2, july 2022 method that is proposed can correctly recognize the face in the picture or video, and it then decides whether or not the subject is wearing a mask [4]. in addition, it can recognize a face even when it is covered by a mask, both while it is moving and while it is being seen on video [21, 22,30], and [23,29]. the approach attains high precision. to properly determine the presence of masks and avoid producing overfitting, we study the ideal parameter values for the convolutional neural network model (cnn) [33-34]. 3. materials and method because lsvm is very fast, we utilized linear svm with non-linear x2-kernel to train the model. to achieve a balance, we use the homogeneous kernel map, which estimates explicit feature mappings to approximate the homogeneous additive kernels, to compute a linear approximation to x2 kernel [32]. the proposed model efficiently resolves the problem and accurately recognizes the faces. in the detection, the process algorithm collects all features of video frames. in the next step, the layering process started where all features are combined and sent for the prediction. in the last step, the prediction step is taken where all features are gathered from in this section, we discuss materials and implementation used for facemask detection on gates. our proposed method consists of four major steps that are: (i) object detection (ii) object localization (iii) objectness + box (iv) classification and confidence score. we first use a detector to extract roi and divide inputs into two parts: the object region, which is crucial, and the context region, which is secondary. the roi score is used to restore the classification outcome. finally, we do classification with a confidence score. the detailed proposed flow is shown in figure 2. figure 2: proposed model for facemask detection and classification. to recognize the faces, we employed the viola-jones algorithm. the three strategies employed by [14] are as follows: cascade, integral image (representation), and adaboost (classifier). several pixel-wise operations are used to compute the integral frame. we used equation 1 to calculate the sum of the left and top pixels of the impacted pixel, and the integral image is calculated. 𝖯(𝑥′, 𝑦′) = ∑ 𝐼(𝑥, 𝑦) 𝑤ℎ𝑒𝑟𝑒𝑥′ ≥ 𝑥𝑎𝑛𝑑𝑦′ ≥ 𝑦 (1) the interest rio locations are found using hessian matrix approximation. rio is determined using equation 3. 𝛽(i, 𝛼) = [ cxx(i, 𝛼) cxy(i, 𝛼)] cxy(i, 𝛼) cyy(i, 𝛼) (2) where cxx(i, 𝛼) , cxy(i, 𝛼) , cyy(i, 𝛼) are gaussian convolution of second order derivative. by forming a rectangle region around the interest locations, the descriptor is extracted. before computing the haar wavelet, the region is partitioned into smaller subregions. we increase the size of the object bounding box by a factor of 1:1 while maintaining the bounding box's center the same. the obtained result illustrates that increasing the bounding box size can increase classification performance. the image and performed the prediction as shown in figure 3. figure 3: model structure of yolo v5[10] the model is exploited with fine-tuned parameters. yolo v5 is used to classify face recognition into with and without masks, as shown in figures 4 [24 26]. the dataset is split into 80% training and 20% testing sets and classes are with masks and without a mask. moreover, the 300 images are from class 1 which is with a mask, and 450 without a mask after that images are augmented using the technique of rotation, translation, and scaling to get more images and overall images become 1200 for with mask and 1350 without a mask. furthermore, we used yolo fine tuned algorithm with change weights as well as activation function. figure 4: facemask detection 4. results and discussion: in this section, we discuss the steps and parameters that were employed during the computation of results. in the facemask detection process, we tested the proposed model. during testing, the proposed model accuracy and error rate were both taken into account while detecting fasemask. the result found are shown in figure 5 and 50 | vol.3 no.2, july 2022 figure 5: confusion matrix table 1 confusion matrix figure 6: fasemask detection results we used the generalized intersection of union (giou) loss function for the yolo v5. the gıou maximizes the overlap area of actual and predicted boundary boxes. as shown in figure 6, you is initially high at the point of prediction then slowly overlaps with the actual value. this shows that you gives good results that achieve good precision which is 97%. the value you are used for training purpose, this graph shows that this loss function maximizes the overlap area of ground truth and predicted boxes as shown in figure 4. these results show that yolo v5 accurately classifies face recognition with and without masks as shown in table 1 and figure 5. in the ground reality, there are 367 faces with masks and 476 without. moreover, yolo v5 is used to detect the face accurately. in figure 6, the objectness plot shows how accurately yolo v5 classifies the faces with and without masks. the objectness loss is due to wrong face recognition in giou prediction. the classification plot shows that the classification loss function deviates from predicting 2 for actual classes and 0 for other classes. this shows that yolo v5 classifies the classes more accurately as shown in the results. furthermore, precision and recall plots show that yolo v5 detects the objects accurately. the high precision shows that yolo v5 accurately predicts the objects. high precision shows the correctness of prediction. the recall plot shows how accurately find all the positives by using the yolo v5 algorithm. the recall is high means our results are good and yolo v5 detects the faces efficiently. the map@0.5 plot shows the mean of average precision at 0.5 you. the map is gradually increasing which shows how accurate the results are and how efficiently the yolo v5 algorithm detects the faces with and without masks. the f1 is a harmonic mean of precision and recall. as shown in the results, precision and recall are high which means our f1 score should also be high which shows the better performance of our used algorithm yolo v5 that classifies the faces with and without masks accurately. our results depict that yolo v5 detects the faces with and with masks more accurately and works efficiently in classification. we also used the public dataset to compare our results to the current state of the art. the outcome is shown in table 2. table 2: state-of-the-art performance evaluation dataset accuracy thermal cheetah dataset 0.942 anki vector robot dataset dataset 0.985 drone gesture control dataset 0.971 egohands dataset 0.982 pascal voc 2012 dataset 0.980 5. conclusion in this paper, the deep learning algorithm yolo v5 is exploited to accurately classify faces with and without masks. yolo v5 used the loss functions giou, objectness, and classification. that shows how accurately yolo v5 detects the objects. the dataset is split into 80% training and 20% testing sets and classes are with masks and without a mask. moreover, the 300 images are from class 1 which is with a mask, and 450 without a mask after that images are augmented using the technique of rotation, translation, and scaling to get more images and overall images become 1200 for with mask and 1350 without a mask. the used algorithm is evaluated through simulations in which the used algorithm yolo v5 precision is 96% and accuracy is 95%. the results show that yolo v5 detects objects more accurately. references [1] a. d. miller, t. b. murdock, and m. m. grotewiel, "addressing academic dishonesty among the highest achievers," theory into practice, vol. 56, pp. 121-128, 2017. [2] li, e. “research on face detection methods”. 4th international conference on signal processing and machine learning, 2021. [3] genemo, m.d. “suspicious activity recognition for monitoring cheating in exams”. proc.indian natl. sci. acad. 88, 1–10 (2022). [4] genemo, m. d. (2022). suspicious activity recognition for monitoring cheating in exams. proceedings of the indian national science academy, 1-10. [5] reddy, s., goel, s., & nijhawan, r. “ real-time face mask detection using machine learning/ deep feature-based classifiers for face mask class n(truth) n(classified) accuracy precision recall f1 score with mask 367 370 95.61% 0.95 0.95 0.95 without mask 476 473 95.61% 0.96 0.96 0.96 mailto:map@0.5 51 | vol.3 no.2, july 2022 recognition”. ieee bombay section signature conference (ibssc), 2021. [6] luh, g. “face detection using a combination of skin color pixel detection and viola-jones face detector” international conference on machine learning and cybernetics, 2014. [7] vidal, a., jha, s., hassler, s., price, t., & busso, c. “face detection and grimace scale prediction of white-furred mice. machine learning with applications”, 8, 100312.2022. [8] alakkari, s., & collins, j. j. “eigenfaces for face detection” a novel study. 12th international conference on machine learning and applications, 2014. [9] jeong, h. j., park, k. s., & ha, y. g. (2018, january). image preprocessing for efficient training of yolo deep learning networks. in 2018 ieee international conference on big data and smart computing (bigcomp) (pp. 635-637). [10] garg, d., goel, p., pandya, s., ganatra, a., & kotecha, k. (2018, november). a deep learning approach for face detection using yolo. in 2018 ieee punecon (pp. 1-4). ieee. [11] al-masni, m. a., al-antari, m. a., park, j. m., gi, g., kim, t. y., rivera, p., ... & kim, t. s. (2018). simultaneous detection and classification of breast masses in digital mammograms via a deep learning yolo-based cad system. computer methods and programs in biomedicine, 157, 85-94. [12] wu, d., lv, s., jiang, m., & song, h. (2020). using channel pruning-based yolo v4 deep learning algorithm for the real-time and accurate detection of apple flowers in natural environments. computers and electronics in agriculture, 178, 105742. [13] george, j., skaria, s., & varun, v. v. (2018, february). using a yolo-based deep learning network for real-time detection and localization of lung nodules from low dose ct scans. in medical imaging 2018: computer-aided diagnosis (vol. 10575, p. 105751i). international society for optics and photonics. [14] magalhães, s. a., castro, l., moreira, g., dos santos, f. n., cunha, m., dias, j., & moreira, a. p. (2021). evaluating the single-shot multibox detector and yolo deep learning models for the detection of tomatoes in a greenhouse. sensors, 21(10), 3569. [15] loey, m., manogaran, g., taha, m. h. n., & khalifa, n. e. m. (2021). fighting against covid-19: a novel deep learning model based on yolo-v2 with resnet-50 for medical face mask detection. sustainable cities and society, 65, 102600. [16] zhuang, zhemin, guobao liu, wanli ding, alex noel joseph raj, shunmin qiu, jingfeng guo, and ye yuan. "cardiac vfm visualization and analysis based on yolo deep learning model and modified 2d continuity equation." computerized medical imaging and graphics 82 (2020): 101732. [17] ardhianto, p., subiakto, r. b. r., lin, c. y., jan, y. k., liau, b. y., tsai, j. y., ... & lung, c. w. (2022). a deep learning method for foot progression angle detection in plantar pressure images. sensors, 22(7), 2786. [18] pouyanfar, s., sadiq, s., yan, y., tian, h., tao, y., reyes, m. p., ... & iyengar, s. s. (2018). a survey on deep learning: algorithms, techniques, and applications. acm computing surveys (csur), 51(5), 1-36. [19] guo, yanming, yu liu, ard oerlemans, songyang lao, song wu, and michael s. lew. "deep learning for visual understanding: a review." neurocomputing 187 (2016): 27-48. [12] zhou, xinyi, wei gong, wenlong fu, and fengtong du. "application of deep learning in object detection." in 2017 ieee/acis 16th international conference on computer and information science (icis), pp. 631-634. ieee, 2017. [21] xiao, youzi, zhiqiang tian, jiachen yu, yinshu zhang, shuai liu, shaoyi du, and xuguang lan. "a review of object detection based on deep learning." multimedia tools and applications 79, no. 33 (2020): 23729-23791. [22] mittal, p., singh, r., & sharma, a. (2020). deep learning-based object detection in low-altitude uav datasets: a survey. image and vision computing, 104, 104046. [23] wu, x., sahoo, d., & hoi, s. c. (2020). recent advances in deep learning for object detection. neurocomputing, 396, 39-64. [24] srivastava, s., narayan, s., & mittal, s. (2021). a survey of deep learning techniques for vehicle detection from uav images. journal of systems architecture, 117, 102152. [25] garcia-garcia, a., orts-escolano, s., oprea, s., villena-martinez, v., martinez-gonzalez, p., & garcia-rodriguez, j. (2018). a survey on deep learning techniques for image and video semantic segmentation. applied soft computing, 70, 41-65. [26] zujovic, j., gandy, l., friedman, s., pardo, b., & pappas, t. n. (2009, october). classifying paintings by artistic genre: an analysis of features & classifiers. in 2009 ieee international workshop on multimedia signal processing (pp. 1-5). ieee. [27] kobylin, o. a., gorokhovatskyi, v. о., tvoroshenko, i. s., & peredrii, o. о. (2020). the application of non-parametric statistics methods in image classifiers based on structural description components. telecommunications and radio engineering, 79(10). [28] dredze, m., gevaryahu, r., & elias-bachrach, a. (2007, august). learning fast classifiers for image spam. in ceas (pp. 2007-487). [29] kobylin, o. a., gorokhovatskyi, v. о., tvoroshenko, i. s., & peredrii, o. о. (2020). the application of non-parametric statistics methods 52 | vol.3 no.2, july 2022 in image classifiers based on structural description components. telecommunications and radio engineering, 79(10). [30] wenger, j., kjellström, h., & triebel, r. (2020, june). non-parametric calibration for classification. in international conference on artificial intelligence and statistics (pp. 178 190). pmlr. [31] p. f. felzenszwalb, r. b. girshick, d. mcallester, and d. ramanan, “object detection with discriminatively trained part based models," ieee tpami , 2009. [32] a. vedaldi and a. zisserman, “efficient additive kernels via explicit feature maps," in cvpr, 2010 [33] wu, z., xiong, y., yu, s., & lin, d. (2018). unsupervised feature learning via non parametric instance-level discrimination. arxiv preprint arxiv:1805.01978. [34] gorokhovatskyi, v. о., tvoroshenko, i. s., & vlasenko, n. v. (2020). using fuzzy clustering in structural methods of image classification. telecommunications and radio engineering, 79(9). [35] chen, r. c. (2019). automatic license plate recognition via sliding-window darknet-yolo deep learning. image and vision computing, 87, 47-56. p-issn : 2715-2448 | e-isssn : 2715-7199 vol.3 no.1 january 2022 buana information tchnology and computer sciences (bit and cs) 28 | vol.3 no.1, january 2022 money check result data management application section verbasar by web-based (case study : perum peruri) topan setiawan 1 study program information systems faculty of computer science, universitas ma’some, bandung, indonesia topansetiawan@masoemuniversity.ac.id nono heryana 2 study program information systems faculty of computer science, universitas singaperbangsa karawang, indonesia nono@unsika.ac.id ‹β› bayu priyatna 3 study program information systems faculty of engineering and computer science, universitas buana perjuangan karawang, indonesia bayu.priyatna@ubpkarawang.ac.id abstrak—pemanfaatan teknologi komputer dan sistem informasi di abad ke-20 ini sudah banyak digunakan, tidak terkecuali pada perusahaan. keuntungan dari penggunaan sistem informasi di perusahaan yaitu dapat menyajikan media pengelolaan data dalam proses bisnis perusahaan di mana sangat diperlukan agar efektifitas dapat tercapai. seksi verbasar merupakan salah satu seksi di perum peruri yang bertugas memeriksa hasil cetakan uang kertas, dengan jumlah karyawan sekitar 100 orang. pengelolaan dan penyimpanan data hasil pemeriksaan uang di seksi verbasar masih terdapat permasalahan. permasalahannya ialah penggunaan microsoft excel sebagai media pengelolaan dan penyimpanan data yang dinilai belum optimal. oleh karena itu dibangun sebuah aplikasi pengelolaan data hasil pemeriksaan uang seksi verbasar berbasis web sebagai pengganti sistem lama agar pengelolaan dan penyimpanan data menjadi lebih optimal dalam proses dan laporannya. pada pembuatan aplikasi ini peneliti menggunakan metode pengambilan data seperti observasi, wawancara, dan studi literatur. sdlc (system development life cylce) model waterfall digunakan untuk metode pengembangan sistem dengan uml (unified modelling language) sebagai modelling tools untuk mengembangkan rancangan sistem informasi. hasil yang diharapkan adalah agar pengelolaan data hasil pemeriksaan uang dapat memberikan sistem yang lebih baik dari sistem lama dalam hal pengelolaan data. kata kunci : aplikasi, sdlc, waterfall, web, pemeriksaan, uang abstract—the use of computer technology and information systems in the 20th century has been widely used, including companies. the advantage of using an information system in the company is that it can provide data management media in its business processes where effectiveness must be achieved. the verbasar section is one of the sections in perum peruri, which is in charge of checking the printouts of banknotes, with around 100 employees. there are still problems in managing and storing data on the results of money checks in the verbasar section. the problem is using microsoft excel as a medium for data management and storage, which is considered not optimal. therefore, a web-based application for data management of the verbasar section money examination results was built as a substitute for the old system so that data management and storage became more optimal in the process and report. in making this application, researchers used data collection methods such as observation, interviews, and literature studies. sdlc (system development life cycle) waterfall model is used for system development method with uml (unified modeling language) as modeling tools to develop information system design. the expected result is that the data management of money check results can provide a better system than the old system in terms of data management.. keywords : application, sdlc, waterfall, web, examination, money i. introduction the rapid and widespread use of computer technology has created a new pattern of life where almost every activity, especially in terms of work, cannot be separated from information systems. the current information system plays a role as a supporter or assistant in time efficiency and working methods. many companies are already using information systems to support all production activities that run in the company [1]. an information system is an organized system that functions in managing information where the information can be useful and has the intent and purpose so that the mailto:si17.iissetiani@mhs.ubpkarawang.ac.id 29 | vol.3 no.1, january 2022 information conveyed can be accepted and achieved [2]. one real example of the use of information systems is in the company. in companies, the use of information systems has become used at this time as a supporting medium for the running of business processes. [3]. verbasar section is one of the sections in perum peruri which is in charge of verifying or checking the printed money [4]. the problem in the big sheet verification section (verbasar) is that the data management of money check results is still using microsoft excel (ms. excel), which is considered ineffective [5]. data on the results of money checks are data on the amount of money (quantity) that each employee and data have checked on cash damage found during the inspection. the use of ms. excel as a data management media is more at risk of being changed or even deleted, either intentionally or unintentionally, because anyone can open the data. in addition, the large number of files stored on the computer makes searching for data difficult. then in terms of the admin's ability to operate ms. excel, which is still inadequate, it makes data management more difficult so that the data report on the results of money checks is hampered [6]. considering these problems, an internal information system application is needed in the verbasar section to manage money audit data so that data management becomes more effective and reduces the number of files stored. therefore, this research takes the title "web-based application for data management of verbasar section money check results." ii. method a. data collection techniques data collection techniques carried out in this study are as follows. 1. observation observation is to make direct observations in the verbasar verification section as a place of research. this observation is intended to obtain information related to the problems that are the object of research. 2. interview at the interview stage, the researcher conducted interviews with the leadership and admin of the verbasar section, regarding the data management of the ongoing money check results. the purpose of this interview is to obtain factual information. 3. literature study this literature study carries out activities of collecting research bases from journals, books, final assignments, proceedings that have a relationship with the research title appointed. information obtained from various sources, both national and international. b. system development method the system development method used in this research is the waterfall model sdlc (system development life cycle). the sdlc method aims to produce a quality system that is by the customer's wishes [7]. the waterfall is analogous to a waterfall, and water will flow step by step until it finally reaches its destination [8]. sdlc waterfall is a software development method that proposes a systematic and sequential approach to software starting from the level of system progress starting from analysis and ending with maintenance [9]. figure 1 stages of the sdlc waterfall model [10] the steps involved in making the application for data management results from money checks are as follows: 1. analysis (analysis) at this stage collect information about the system that is running in the verbasar section by means of observation and interviews. observations and interviews aim to obtain data that is in accordance with the facts in the verbasar section. after finding the problem, then it is analyzed to find a solution in the form of a proposed system. 2. design this stage is done before coding. this stage has 2 parts, namely system design and interface design. the system design is done with uml tools using 5 diagrams, namely use case diagrams, activity diagrams, sequence diagrams, and class diagrams. the second part is interface design, so that the interface is more effective, which means it is ready to be used with the desired results. the needs in question are the needs of users. the user interface on a system will affect the performance of its users. user interface design using the pencil application. 3. implementation (coding) in this stage, programming is carried out or the process of translating the design of the system design into a form that can be understood by the machine, using programming language codes. making this program or application uses the codeigniter framework, bootstrap as a css framework, mysql as a database, php programming language, and xampp as a web server. 4. testing testing is the stage of feasibility testing on the system to be made, starting from the system design, to the function of each of its features. this test uses black box testing and white box testing methods. 30 | vol.3 no.1, january 2022 5. maintenance maintenance is the final stage in the waterfall model by performing maintenance on the current system. system maintenance can be done by making repairs to parts that experience problems when running and backing up data regularly. backing up data can take advantage of external storage such as a hard disk or by storing it in cloud storage (cloud storage). iii. results and discussion a. system planning at this stage is an activity to identify, analyze and evaluate problems or obstacles that occur in the current system, namely the process of managing data from the results of money checks. with the system analysis, it is hoped that improvements to the new system that will be designed can be proposed. 1. current system analysis the following is a flowmap of the system running on the data management of the results of money checks in the verbasar section: figure 2 flow maps of a running system in the system that runs on the process of managing the data on the results of the examination of money, the group leader asks for a form to write the data on the results of the examination of money which consists of 2 forms, namely a form to write the amount of money checks carried out by employees and a form to write the amount of money damage. the inspection result data form is filled in by the group leader, then inputted into microsoft excel according to each group, and the money damage data form is given to the admin. admin recaps data on damage to money and makes a report on the results of the inspection to be submitted to the leadership. leaders receive reports and acc reports. 2. analysis of the proposed system judging from the analysis that has been carried out on the system currently running in the verbasar section, it is proposed to make an application to be used as a medium or a place to manage data on the results of money checks. this application is a substitute for ms. excel as a data manager for money check results. figure 3 flow maps of the proposed system in the flowmaps, the proposed system for the data management application from the results of the money check requires 2 users, namely the admin and the group leader with their respective access rights. the first user or user is the group leader, where the task of the group leader is to input data on the results of money checks carried out by his group members. each group leader will have their own account to be able to access the system or application. the second user is the admin, where the admin has the task of recapitulating the data on money damage obtained from each group leader and making reports to the leadership regarding the data on the results of money checks in the verbasar section. b. system design the system modeling design uses uml (unified modeling language) which is a standard modeling language used as a visualization, specifying, constructing, and documenting the tools of an object-oriented software system [11]. 31 | vol.3 no.1, january 2022 1. use case diagrams use case diagrams only describe a condition seen by outside users, not how the functions that exist in the system [12]. the use case diagram for the application for managing data on the results of money checks consists of the admin and group leader. a. admin: managing employee data, checking group data, managing user data, managing group data, managing goods data, managing damage type data, managing damage data, and managing reports. b. group leader: the group leader is an actor who has access rights to manage data on the results of money checks that have been carried out by employees who are members of the group. figure 4 use case diagrams 2. activity diagrams activity diagram is a series of processes used to describe activities that are formed in a series of operations. activity diagrams have the function of showing the sequence of process activities on the system, helping to understand the overall process [13]. the activity diagram below is an explanation of the process of managing money audit data which consists of adding, deleting data, and filtering data on money inspection results by the group leader. figure 5 activity diagram managing data from money check results 3. sequence diagram sequence diagram is a diagram of the interaction between objects based on a certain time sequence. sequence diagrams start from pulling the actors in the use case diagram by creating a detailed sequence of the flow of the use case process with messages flowing in it [14]. in the sequence diagram, the data input of the results of the money check carried out by the employee is entered by the group leader into the application, and the data will be stored in the production table in the database. gambar 1 sequence diagram input data on money check results 4. class diagrams a class diagram is a uml (unified modeling language) diagram that defines the classes and the relationships between the classes to be built. the classes in the class diagram consist of attributes and methods to make creating programs that describe the relationship between design documentation and appropriate software [15]. the class diagram on the application for managing data on the results of money checks can be seen in the following figure. 32 | vol.3 no.1, january 2022 figure 7 class diagrams 5. interface design interface design is an important part that will interact directly with users. the interface design that will be built on the money audit result data management application is as follows. a. main page display design with admin access: figure 8 admin access main page b. main page display design with access as group leader: figure 9 main page access group leader c. money check results page done by employee: figure 10 data page of money check results d. damage data page which is the result of findings on money checks carried out by employees: figure 11 money damage data page c. system implementation system coding is the translation stage from the designer that has been made in the uml model and the system interface design. at this stage the application is made using the codeighniter framework in which there is a php programming language combined with html, css, boostrap and javascript, for the base using mysql. 1. login page the login page is used by the user to enter the market retribution application according to the access rights owned by the user. figure 2 login page 2. dashboard page the dashboard page is the main page after the user has successfully logged in. 33 | vol.3 no.1, january 2022 figure 13 admin dashboard page figure 14 group leader dashboard page 3. money check result page this page contains data on the number of cash checks completed by employees. figure 15 officer access page 4. money damage data page this page will display data on cash damage collected from the results of money checks by employees. figure 16 money damage data page d. system testing system testing on the money check result data management application uses the black box method where the test is carried out on the program display to ensure the program runs well as desired. the testing process for this application is as follows. table 1 testing the login page process design output result open the app show login page matching fill in username and password with correct data a successful login message appears and the application dashboard appears according to the user matching fill in username and or password with wrong data dashboard fails to appear, a message appears that the username or password has not been registered matching empty the username and or password dashboard fails to appear, a prompt message appears to fill in the username and password form matching e. maintenance maintenance on the system is carried out by correcting the program in case of errors, adding features as needed, and regularly backing up data. how to backup data by downloading the report on the results of the money check, then creating a report folder with the name of the date and month of the report. in addition, data related to the system, if there is a system update, data backups must be carried out, so that the data stored is the latest data. backup data storage can use 2 ways, namely with external storage such as hard disks and cloud storage (cloud storage). cloud storage is intended to avoid data loss when the hard drive is lost or damaged. system maintenance aims to maintain or even develop the system that has been created. iv. conclusion based on the descriptions in the previous chapters, a problem was found in managing the data on the results of money checks in the verbasar section, which still uses ms. excel as a data management and data storage tool. this problem is the reason for creating a web-based money check result data management application. from the results of trials that have been carried out on the application, the following conclusions can be drawn: 1. this application can perform data processing on the results of money checks, including adding, changing, and deleting data. 34 | vol.3 no.1, january 2022 2. this application can display data reports on the results of money checks, which can be used to report to the leadership. 3. the data stored in this application is more efficient because the data is collected in one web-based storage source. 4. reports on the results of money checks can be displayed based on a certain period with the provided data filter facility. reference [1] suhendro, dedi. 2017. “perancangan dan implementasi realisasi anggaran pendapatan ( studi kasus : pengadilan negeri klas ib pematangsiantar ).” seminar nasional teknologi informatika, 30–36. [2] a. lia hananto et al., “analysis of drug data mining with clustering technique using k-means algorithm,” j. phys. conf. ser., vol. 1908, no. 1, 2021. [3] h. bagir and b. e. putro, “analisis perancangan sistem informasi pergudangan di cv. karya nugraha,” j. media tek. dan sist. ind., vol. 2, no. 1, p. 30, 2018. [4] n. nurhasanah and p. inoprasetya, “kompetensi, komitmen organisasi, dan motivasi terhadap kinerja karyawan verbasar perum peruri karawang,” balanc. econ. business, manag. account. j., vol. 18, no. 1, p. 77, 2021. [5] h. mulyono, “124-335-1-pb (1),” vol. 2, no. 4, pp. 771–780, 2017. [6] merri parida, s.kom, and williams kurnia wardany. 2019. “sistem informasi pengolahan data produksi brebasis web pada cv semangat jaya lampung.” hilos tensados 1: 1–476. [7] a. h. hendrawan, “rancang bangun sistem informasi hasil produksi dengan menerapkan metode system development life cycle,” ranc. bangun sist. inf. has. produksi dengan menerapkan metod. syst. dev. life cycle, pp. 1–7, 2016. [8] mujilan, agustinus. 2017. analisis dan perancangan sistem. madiun: fakultas ekonomi dan bisnis universitas katolik widya mandala. [9] huda, baenil, and saepul aripiyanto. 2019. “berbasis android dan web monitoring (penelitian dilakukan di kab. karawang).” teknologi informasi 4 (1): 11–24. [10] duha, n. j., suryadi, s., yanris, g. j., simanjuntak, n. j., suryadi, s., silaen, g. j. y., manajemen, a., & komputerlabuhan, i. (2017). sistem pengarsipan surat bagian organisasi dan tatalaksana. 5(3), 26–36. [11] d. w. t. putra and r. andriani, “unified modelling language (uml) dalam perancangan sistem informasi permohonan pembayaran restitusi sppd,” j. teknoif, vol. 7, no. 1, p. 32, 2019. [12] s. rosyida and v. riyanto, “sistem informasi pengelolaan data laundry pada rumah laundry bekasi,” jitk (jurnal ilmu pengetah. dan teknol. komputer), vol. 5, no. 1, pp. 29–36, 2019. [13] e. erlinda, “pengolahan data sensus penduduk menggunakan bahasa pemrograman php berbasis web pada kecamatan bukit sundi kabupaten solok,” j. teknol. dan open source, vol. 1, no. 1, pp. 46–57, 2018. [14] m. i. hanafri, triono, and i. luthfiudin, “rancang bangun sistem monitoring kehadiran dosen berbasis web pada stmik bina sarana global,” j. sisfotek glob., vol. vol.8, no. no.1, pp. 81–86, 2018. [15] sidik, achmad, edy tekat, bronto waluyo, and siti susilawati. 2018. “perancangan sistem informasi manajemen produksi di pt aneka paperindo sejahtera.” jurnal sisfotek global 8 (2): 8–13. p-issn : 2715-2448 | e-isssn : 2715-7199 vol.3 no.1 january 2022 buana information tchnology and computer sciences (bit and cs) 5 | vol.3 no.1 january 2022 application of recapitulation and staff performance assessment using standard working method agustia hananto 1 study program information system faculty of engineering and computer science, universitas buana perjuangan karawang agustia.hananto@ubpkarawang.ac.id eko pramono 2 study program computer science faculty of engineering and informatics bina sarana informatika university indonesia eko.eop@bsi.ac.id ‹β› baenil huda 3 study program information system faculty of engineering and computer science, universitas buana perjuangan karawang baenil88@ubpkarawang.ac.id abstrak—evaluasi atau penilaian kinerja karyawan diperlukan untuk menilai pencapaian kerja atau prestasi kerja dari karyawan perusahaan terhadap target yang diberikan dalam waktu tertentu. tujuan penelitian adalah untuk menganalisis metode rekapitulasi dan penilaian kinerja staf serta untuk membuat model rancang bangun aplikasi rekapitulasi dan penilaian kinerja staf. metode penelitian yang digunakan adalah penelitian kualitatif dimana pengumpulan data dilakukan dengan cara observasi dan wawancara. pengembangan aplikasi berdasarkan metode iterative waterfall digunakan untuk mengakomodir proses rekapitulasi dan penilaian kinerja karyawan tersebut. manfaat dari penelitian yang dilakukan selain untuk mengetahui dampak dari keterlambatan informasi, juga untuk mempercepat proses rekapitulasi dan penilaian kinerja itu sendiri. informasi mengenai pencapaian kerja setiap bulannya disajikan dalam aplikasi yang dapat dilihat oleh masing-masing karyawan. hasil akhir rekapitulasi dan penilaian berupa nilai total dan rasio pencapaian prestasi kerja diharapkan mampu mempermudah kepala bagian dalam mengambil keputusan. kata kunci: sistem informasi, karyawan, rekapitulasi, peniliaian kinerja abstract—employee performance evaluation or appraisal is needed to assess company employees' work achievement or work performance against the given target within a specific time. the purpose of the study was to analyze the method of recapitulation and staff performance appraisal and create a design model for the application of recapitulation and staff performance appraisal. the research method used is qualitative research, where data collection is done using observation and interviews. application development based on the iterative waterfall method accommodates the recapitulation process and employee performance appraisal. the benefits of the research carried out are to determine the impact of information delays and speed up the recapitulation process and the performance appraisal itself. every month, information regarding work achievements is presented in an application that each employee can see. the final results of the recapitulation and assessment in the form of total values and ratios of work achievement are expected to facilitate the head of the department in making decisions. keywords: information systems, employees, recapitulation, performance appraisal i. introduction evaluation and assessment is a thing that is commonly found in an organization, both academic, social and corporate organizations. performance evaluation or assessment to individuals, work units, or work sub-units is closely related to awarding awards set by the institution or organization. performance or performance can be defined as an individual or teamwork following predetermined work instructions. evaluation or assessment of work performance is a routine program carried out by an organization, both government and private agencies, to determine steps in fostering or developing employees or employees following the results obtained from the evaluation or assessment. evaluation or assessment of work performance is a routine program carried out by an organization, both government and private agencies, to determine steps in fostering or developing employees or employees following the results obtained from the evaluation or assessment [1]. based on the interview results, following the company's internal policy that the recapitulation of work results is one of the elements in the performance appraisal. the assessment results will be submitted to each staff after the semester assessment period ends. in this case, the staff cannot improve because the final score results are final, and the team can only enhance work achievement in the following semester. this happens because there is no information to staff regarding work achievements every month in the current semester period. submission of information on work achievement every month to staff can have an indirect effect on the performance of the team themselves [2]. the results of interviews that have been conducted show that information about work achievements in the previous month is essential for the staff themselves. as a result, if the work achieved during the last month does not reach the target or is not good, the team can improve the following month. an integrated information system or application is expected to accommodate these needs by providing output in monthly work achievement data. implementing the correct application can simplify and speed up evaluating employee performance [3]. so that feedback can be informed quickly to make the necessary 6 | vol.3 no.1, january 2022 improvements. so it is hoped that the built application can help optimize the performance of company employees [4] . and facilitate the supervision process [5]. definition of performance evaluation performance appraisal is a formal management system provided for evaluating the quality of individual performance in an organization [1]. staff definition staff are people who help a leader or chairman in managing something or a certain work unit [1]. for example, office staff, administrative staff, management staff, and so on definition of standard working method according to mondy (2008), the working standards method is a method of evaluating performance by comparing work achievements with predetermined targets or expected outputs. standards reflect the normal output of an average employee working at a normal pace. this assessment method can be applied in almost all types of work, but is generally used in production lines[6]. define flow map flow map according to wahyudi (2012)[7], "a diagram that shows the flow of data in the form of forms or information in the form of documentation that flows or circulates in a system. this diagram serves to determine the relationship between entities in a system. definition of unified modeling language according to sri mulyani (2016), unified modeling language (uml) is a standard tool for a system development technique that uses a graphical language as a tool for documenting and performing system specifications [8]. definition of php and codeigniter php stands for recursive php: hypertext preprocessor. php is a programming language that is specifically used for web application development [9]. codeigniter is an open-source php framework or framework that can help speed up developers in creating or developing a web-based application [10]. ii. method a. data collection techniques the research was conducted using qualitative methods, and data collection techniques were carried out by field observations, interviews, and literature studies [11]. 1. observation 2. the observation step was carried out by making direct observations of the operational contact center of pt. xyz to get information or information that is directly related to the problem. 3. interview interviews were conducted by giving several open and closed questions to the parties involved in the operational activities of the organization, namely department heads, team leaders, and staff. 4. literature study books, literature, or library materials are used in research by taking notes and or citing expert opinions as supporting the theoretical basis of research. b. system development method the system development method used in this research is iterative waterfall. iterative waterfall has advantages compared to the classic waterfall, namely it has feedback at each stage to the previous stage [12]. the choice of this method is not only because it is easy but also has advantages, namely when the system requirements are fully and explicitly defined, the system development will run well [13]. the steps or stages of system development with the iterative waterfall method are as follows [12]: 1. analyze analyze is the stage where all things related to system development are analyzed. so that problems and needs are identified as well as solutions that can be applied.. 2. design the results of the analysis carried out, the researchers then made an overview of the current process flow and the needs needed to make a design. the design made includes the application design using the unified modeling language (uml) as the model. 3. coding the coding stage is the stage of pouring the results of the design into codes in programming languages [12]. the programming language used is php. the coding process follows the general rules and rules used in the code igniter framework[10]. 4. testing the process of testing the application from the coding results is done by means of black box testing. where the functions of the application such as links, buttons, forms, and other elements can work or not [14]. 5. implementation the implementation stage is carried out by researchers by installing and configuring applications on the server computer. in addition, socialization is also carried out to all users or users who have access to the application. 6. maintenance maintenance actions are in the form of correcting errors found during testing and when the application is running. in addition, maintenance is also carried out on the application and database, namely cleaning the cache on the server computer and periodically backing up the database. figure 1 iterative waterfall[12] iii. results and discussion a. current system analysis the research was conducted at the contact center of pt. xyz. contact center pt. xyx has a role in providing aftersales service to consumers, as an information service center 7 | vol.3 no.1, january 2022 for consumers via telephone, short message or short message service (sms), whatsapp, electronic mail, and social networking media. from the results of observations and interviews, it was found that the performance appraisal of staff working at the contact center of pt. xyz consists of four elements, namely work productivity, consumer satisfaction index or cs index, attendance, and product knowledge. the assessment process is carried out every six months or every semester. the staff team leader will recapitulate all elements of the assessment and perform calculations according to the calculation reference that has been determined. the results of the recapitulation and calculations will be submitted to the head of the department. after the assessment data is sent to top management, the department head will inform each staff of the performance value. the performance value is the final value, so if there is an element of an unfavorable assessment, it cannot be improved. staff can make improvements in the next assessment period. figure 2 current state flow map calculation method by determining the achievement of each element included in the range of certain classes in accordance with the targets and calculation provisions that have been determined. the result of the calculation with the target is expressed as a percentage. furthermore, the percentage of the comparison results will be multiplied by the weight that has been set as well. the accumulated results of the weighting of all elements in the form of a percentage number are expressed as the final value or key performance indicator (kpi). the formula for calculating the final value is stated as follows: 𝐅𝐢𝐧𝐚𝐥 𝐬𝐜𝐨𝐫𝐞 ∑(𝐖𝐨𝐫𝐤 𝐚𝐜𝐡𝐢𝐞𝐯𝐞𝐦𝐞𝐧𝐭 𝐜𝐥𝐚𝐬𝐬𝐢𝐟𝐢𝐜𝐚𝐭𝐢𝐨𝐧 𝐱 𝐖𝐞𝐢𝐠𝐡𝐭) description final score : kpi, expressed in % work achievement classification : work achievement is compared with the calculation reference set by the department. expressed in % weight : determined by internal department, expressed in % table 1 classification of work achievements job achievement classification/criteria (%) x1 ~ x2 xx % x2 ~ x3 xx % x3 – x4 xx % ... ... xn-1 ~ xn xx % b. system design 1. system proposal referring to the current conditions in the recapitulation process and performance appraisal of the contact center staff of pt. xyz, the researcher proposes to build a web-based application to facilitate the process of recapitulation and staff performance appraisal, so that staff can see their respective work achievements every month and also as a monitoring support tool for department heads. the flow map of the application submitted is as follows: figure 3 the proposed system flow map in the flow map of the proposed system, staff or agents can see the achievement of work every month. so that if there are work achievements in the previous month that are not good, improvements can be made during the same assessment period. 2. system requirements the requirements in the system are outlined in the following table: table 2 system requirements function or position requirement accees dept.head agent 8 | vol.3 no.1, january 2022 team leader manage productivity data, cs index, attendance, and product knowledge test scores. save the assessment standards that have been set. admin staf view productivity data, cs index, attendance, and product knowledge test scores and their respective final scores agent head of department view productivity data, cs index, absenteeism, and product knowledge test scores, and the final score of all staff. save the assessment standards that have been set. department head by using the unified modeling language (uml), the access requirements in the system are described in the following use case diagram:: figure 4 use case diagram of the proposed system 3. activity diagram activity diagrams represent various activities or pieces of processes and sequences in the system [12]. figure 5 activity diagram calculate the final score (kpi) 4. sequence diagram sequence diagrams are used to model the interactions between objects and also represent sample snippets of software system processes [15]. figure 6 sequence diagram calculate the final value 5. class diagram class diagrams describe the static structure of a system or how the system is structured. the static structure of a system consists of a number of classes and their dependencies. a class represents an entity that has attributes and functions[12]. figure 7 the proposed system class diagram 6. database design database or database is needed to store data that will be, is being, or has been processed in the application. the relationship between tables in a database consisting of several entities is as follows. 9 | vol.3 no.1, january 2022 figure 8 relationships between tables in the database 7. interface design layout design describes the interface page (interface) designed on the application. figure 9 application login page display design after successful login, on the main page or dashboard, will display the user profile. the menu list is on the left. figure 10 dashboard page display designs the following is the final value calculation page display design. figure 11 staff kpi calculation page display design c. system implementation the coding process or coding to translate the design that has been made into a program using the codeigniter php framework which is paired with the adminlte3 site. the application built is called logsheet. figure 12 login page display figure 13 dashboard page view 10 | vol.3 no.1, january 2022 figure 14 kpi calculation page view by staff a. system test system testing with blackbox testing is carried out to ensure that the functions in the applications built do not contain errors or bugs when the application is run. table 3 testing the final value calculation menu (kpi) process design description result select the "kpi by agent" menu display the final score calculation page or kpi by staff name matching in kpi by agent, choose the name of the staff, the period of the first month and the end of the month displays a selection of staff name, starting month and ending month. login agent, unable to select staff name matching in kpi by agent, choose submit "go" after selecting staff name, starting and ending month period displays a summary of work achievements and final grades (kpi) based on the selected staff and month matching select the "kpi assessment" menu displays the kpi assessment page which contains a summary of the kpi achievements of all staff matching iv. conclusion some conclusions from the research this: 1. the developed web-based application can simplify evaluating staff performance at the contact center of pt. xyz. the calculation method in the application follows the provisions or standards that exist in the internal organization, both the targets to be achieved and the assessment reference. 2. the application developed can help staff monitor work achievement every month. so that if there are poor work achievements in the previous month, staff can find out more quickly and make performance improvements in the following month. in addition, the application can also be helpful for department heads if a special evaluation is needed for staff based on final grades. reference [1] m. abdullah, manajemen dan evaluasi kinerja karyawan. sleman: aswaja pressindo, 2014. [2] d. a. wardani, “pengaruh penerapan aplikasi sistem informasi akuntansi terhadap kinerja karyawan pada pd. bpr rokan hulu pasir penga iran,” j. chem. inf. model., vol. 53, no. 9, pp. 1689–1699, 2017. [3] y. suherman and d. yadewani, “aplikasi sistem informasi penilaian kinerja karyawan,” j-click, vol. 6, no. 2, pp. 201– 207, 2019. [4] i. m. hadi, t. tukino, and a. fauzi, “sistem informasi monitoring evaluasi standar pembelajaran menggunakan framework codeigniter,” ciastech 2020, no. ciastech, pp. 443–452, 2020. [5] v. felita, k. saputra, s. keputusan, k. pendidikan, and b. lampung, “aplikasi monitoring kerja karyawan ( e-kinerja ) berbasis web menggunakan framework codeigniter di citra angkasa tercipta ( cat ) bandar lampung,” j. vania, pp. 1–8, 2020. [6] a. marbawi, manajemen sumber daya manusia. lhokseumawe: unimal press, 2016. [7] z. makmur, “pengembangan sistem informasi permintaan pembelian kebutuhan kantor pada dealer management system,” j. teknosain, vol. xv, no. 3, pp. 78–88, 2018. [8] k. yuliana, saryani;, and n. azizah, “percanangan rekapitulasi pengiriman barang berbasis web,” j. sisfotek glob., vol. 9, no. 1, 2019, doi: http://dx.doi.org/10.38101/sisfotek.v9i1.223. [9] r. sabaruddin and w. e. jayanti, jago ngoding pemrograman web dengan php, no. january. surabaya: cv. kanaka media, 2019. [10] i. daqiqil, framework codeigniter sebuah panduan dan best practice. pekanbaru, 2011. [11] b. huda and s. aripiyanto, “aplikasi sistem informasi lowongan pekerjaan berbasis android dan web monitoring (penelitian dilakukan di kab. karawang) 1baenil,” j. buana ilmu, vol. 4, no. 1, pp. 11–24, 2019, doi: https://doi.org/10.35706/sys.v1i2.2076. [12] r. mall, fundamentals of software engineering fourth edition, 4th ed. delhi: phi learning private limited, 2014. [13] b. huda and b. priyatna, “penggunaan aplikasi content management system (cms) untuk pengembangan bisnis berbasis e-commerce,” systematics, vol. 1, no. 2, pp. 81–88, 2019, doi: https://doi.org/10.35706/sys.v1i2.2076. [14] g. maulani, d. septiani, and p. n. f. sahara, “rancang bangun sistem informasi inventory fasilitas maintenance pada pt. pln (persero) tangerang,” icit j., vol. 4, no. 2, pp. 156–167, 2018, doi: 10.33050/icit.v4i2.90. [15] b. rumpe, modeling with uml. aachen: springer, 2016. p-issn : 2715-2448 | e-isssn : 2715-7199 vol.3 no.2 july 2022 buana information tchnology and computer sciences (bit and cs) 59 | vol.3 no.2, july 2022 analysis of signal quality, voice service, and data access on telkomsel and indosat providers in pakisjaya district arip solehudin 1 study program technical information universitas singaperbangsa karawang arip.solehudin@gmail.com nono heryana 2 study program information system universitas singaperbangsa karawang nono@unsika.ac.id ‹β› agustia hananto 3 study program information system universitas buana perjuangan karawang agustia.hananto@ubpkarawang.ac.id abstract—kualitas dalam komunikasi sangat berpengaruh dalam kegiatan di zaman modern ini. kesenjangan pada kualitas jaringan yang diharapkan terhadap kualitas jaringan sebenarnya tak jarang ditemui. pengukuran terhadap performa jaringan penting untuk mereferensikan kepada pengguna untuk memilih operator yang tepat sekaligus bagi penyedia layanan untuk memperbaiki kualitas pelayanannya metode drive test merupakan metode yang digunakan untuk mengukur performa jaringan seluler. tujuan penelitian ini yaitu untuk mengetahui hasil perbandingan performa antar provider penyedia layanan. parameter yang digunakan dalam analisa ini adalah receive signal code power (rscp), cell id, signal noise ratio (snr), upload rate, dan download rate. parameter ini digunakan untuk proses pengukuran metode drive test. sedangakan alat bantu yang digunakan untuk pengumpulan nilai dari masing-masing parameter drive test dengan menggunakan aplikasi berbasis android, gnet track pro. hasil pengukuran dengan gnet track pro ini kemudian diolah dengan aplikasi google earth sebagai pemetaan area pengukuran.. data penelitian yang dikumpulkan dalam analisa ini sebanyak 60 kasus, dengan setiap desa berjumlah 20 data untuk selanjutnya diolah dengan metode drive test. dalam penelitian ini, telkomsel cenderung unggul dalam layanan voice dengan 150 jumlah panggilan berhasil dari 150 percobaan panggilan. sementara indosat unggul dalam penyediaan layanan akses data dan sinyal terutama di desa tanjungbungin. kata kunci: drive test, gnet track pro, telkomsel, indosat abstract—quality in communication is very important in activities in this modern era. it is expected that the expected network quality against network quality is actually not expected. measurement of network performance is important to refer for users to choose the right operator for service providers to improve the quality of their service drive test method is a method used to measure cellular network performance. the purpose of this study is to study the results of research between service providers. the parameters used in this analysis are receive signal code power (rscp), cell id, signal noise ratio (snr), upload rate, and download rate. this parameter is used for the measurement process of the drive test method. while the tool used to collect the value of each drive test parameter by using an android-based application, gnet track pro. the measurement results with gnet track pro are then processed with the google earth application as a measurement area. research data collected in this analysis were 60 cases, with each village 20 data collected to continue processing with the drive test method. in this study, telkomsel chose to excel in voice services with 150 successful numbers of 150 attempted calls. while indosat excels in providing data access services and special signals in the village of tanjungbungin. keywords: drive test, gnet track pro, telkomsel, indosat i. introduction communication is an integral part of our social life. communication is very closely related to support socialization. without communication, civilization and socialization would not run as it is today. in today's era, the ease of communication is one of the advantages of technological advances. unlike in the past where to communicate over long distances, humans still rely on postal mail, which takes days to be received by the recipient of the letter, technology makes communication easier, more practical, and can even be done in real-time. usually, operator services in densely populated areas such as urban areas tend to be better than areas with low population density, such as in rural areas. this is actually influenced by many factors, including the amount of network infrastructure and optimal operational capacity placed in an area, to the number of network users in that area. this certainly affects the user experience in communicating, poor signal conditions will have an impact on dropped calls, unstable data connections, and ultimately affect the user's decision to choose a network operator that suits their needs. in theory, the density and range of the signal will affect the speed of the data access connection itself. lack of coverage area has the potential to reduce the average data access speed in an area and vice versa. however, this could be due to two factors. the first is the result of data collection methods that rely on the quantity of user samples in an area to cause a lack of data in an area, to the quality of the sample itself, which may represent the actual measurement results in that area. it is hoped that this research can provide real aspects of the telkomsel and indosat provider networks so as to provide an answer to the actual situation of the network where in the initial research conducted by the author the data produced is quite contradictory, becomes a guideline and preference for the community in determining the use of the right operator. 60 | vol.3 no.2, july 2022 ii. study of literature 2.1 mobile network development mobile wireless communication or wireless telecommunications has evolved in a short span of time, and as it should be, the impact of these innovations is felt by all walks of life significantly. as humans, we are designed to seek wider and deeper connections, and mobile technology has opened up possibilities for us to communicate better and more easily. the development of cellular technology started with 1g and was followed by 2g, 3g, 4g and 5g technology is currently being developed. 2.2 quality of service (qos) according to [1], quality of service (qos) is a method of measuring how good the network is and is an attempt to define the characteristics and properties of a service. qos is used to measure a set of performance attributes that have been specified and associated with a service [6]. in general, quality of service (qos) is a measurement method used to determine the capabilities of a network such as; network applications, hosts or routers with the aim of providing better and planned network services so that they can meet the needs of a service. through qos a network administrator can give priority to certain traffic. qos offers the ability to define the attributes of the services provided, both qualitatively and quantitatively. the purpose of qos is to provide different quality of service based on service needs in the network. according to [2] quality of service (qos) is a method of measuring how good the network is and is an attempt to define the characteristics and properties of a service. qos is used to measure a set of performance attributes that have been specified and associated with a service. in general, quality of service (qos) is a measurement method used to determine the capabilities of a network such as; application network, host or router with the aim of providing a better and planned network service so that it can meet the needs of a service. there are several parameters in qos, namely; bandwidth, throughput, jitter, packet loss, and latency.2.1 mobile network development. mobile wireless communication or wireless telecommunications has evolved in a short span of time, and as it should be, the impact of these innovations is felt by all walks of life significantly. as humans, we are designed to seek wider and deeper connections, and mobile technology has opened up possibilities for us to communicate better and more easily. the development of cellular technology started with 1g and was followed by 2g, 3g, 4g and 5g technology is currently being developed. 2.3 mobile kpi standard the mobile kpi standard consists of 3 aspects. accessibility is the user's ability to obtain services in accordance with the services provided by the network provider. retainability is the ability of users and network systems to maintain services after the service has been obtained until the time limit for the service is stopped by the user. while integrity is the degree of measurement when the service is successfully obtained by the user.[3] 2.4 drive test one method of measuring network data is the drive test. [4] argues that: the drive test is one of the steps for cell planning for a cellular phone network. the test drive retrieves the information needed for cell development planning or cell optimization to create quality communications. the special devices used in the test drive are generally large, separate (laptop, gps, and handphone) and are fewer simple devices. 2.5 gnet track pro according to [5] g-net track pro is an android-based application to perform netmonitoring of umts/gsm/lte/ cdma/evdo networks. this application monitors services from cellid, level, qual, mcc, mnc,lac, adjacent cell service cell time and level. in addition, this application can also be used to determine the quality of voice services with voice sequences, data services with sequence data and test data, and sms services with sms sequences. figure 1. gnet track pro interface display this g-net track pro application can be used to carry out indoor and outdoor test drives and retrieve and visualize data from cells that will be taken on a map or map, the visualization will be presented in the form of a route that has been traversed. the map is marked with an indicator in the form of color and cell breathing and the user will appear on the map. the results of the test drive will be saved in .kml format and a text file that can be extracted on a google map. mechanism in doing this test drive that is by first installing the g-net track pro software on the smartphone that will be used to do a test drive. 61 | vol.3 no.2, july 2022 iii. results and discussion 3.1 drive test analysis using an android smartphone and gnet track pro software, a simultaneous drive test was carried out in 3 villages that were the object of research (tanjungbungin, solokan, and teluk buyung). after that, the data from the test drive is processed in the form of mapping (google earth) and cumulative data in tabular format. in the mapping image, what appears in colored dots is the area traversed during the drive test process. figure 3. drive test process map on gnet track pro table 1. results of cumulative drive test data processing the data that is processed from the results of the drive test for signal parameters is rscp. for internet data, the parameters measured are upload rate, download rate, and ping rate. meanwhile for the voice stage, the parameters measured are call setup, successful call, and dropped calls. then to determine the performance of each parameter, measurements were made with qos standards and kpi standards. tabel 2. standardization of rscp signal sequence parameter performance based on kpi standard tabel 3. standardization of upload rate parameter data sequence performance based on kpi standard tabel 4. standardization of data sequence parameter download rate performance based on kpi standard tabel 5. standardization of ping rate parameter data sequence performance based on kpi standard tabel 6. standardization of voice sequence parameter ping rate performance based on kpi standard 3.2 discussion of signal sequence in the signal sequence, the parameter being tested is rscp. signal strength performance with the rscp kpi standard of telkomsel has a very good predicate in tanjungbungin village, while in solokan village, telkomsel and indosat have good and bad marks, respectively. in telukbuyung village, telkomsel and indosat have bad and good predicate respectively. 3.3 discussion of sequence data in the data sequence, the parameters tested are download rate, upload rate, ping rate. for download rate performance with standard kpi throughput, tiphon and gnet track pro, telkomsel and indosat each have a very bad rating in tanjungbungin village, while in solokan village, location tanjungbungin solokan telukbuyung operator telkomsel indosat telkomsel indosat telkomsel indosat parameter signal rscp -77 -79 -88 -90 -126 -85 db data upload rate 354 544 252 226 324 298 kb/s download rate 20 20 21 14 37 10 kb/s ping rate 407 180 913 238 82 2002 ms voice call setup 50 times successful call 50 50 50 49 50 42 times dropped calls 0 0 0 1 0 8 times location tanjungbungin solokan telukbuyung operator telkomsel indosat telkomsel indosat telkomsel indosat parameter signal rscp -77 -79 -88 -90 -126 -85 decibel result very good very good good not good bad good location tanjungbungin solokan telukbuyung operator telkomsel indosat telkomsel indosat telkomsel indosat parameter data upload rate 354 544 252 226 324 298 kbps result not good not good bad bad bad bad location tanjungbungin solokan telukbuyung operator telkomsel indosat telkomsel indosat telkomsel indosat parameter data download rate 20 20 21 14 37 10 kbps resul t very bad very bad very bad very bad very bad very bad location tanjungbungin solokan telukbuyung operator telkomsel indosat telkomsel indosat telkomsel indosat parameter data ping rate 407 180 913 238 82 2002 ms result bad enough bad bad good bad location tanjungbungin solokan telukbuyung operator telkomsel indosat telkomsel indosat telkomsel indosat parameter voice call attempt 50 times successful call 50 50 50 49 50 42 times dropped calls 0 0 0 1 0 8 times result cssr 100% 100% 100% 98% 100% 84% figure 2. drive test process on gnet track pro 62 | vol.3 no.2, july 2022 telkomsel and indosat each have a very bad rating. in telukbuyung village, telkomsel and indosat each have a very bad predicate. in terms of upload rate performance with standard kpi throughput, tiphon and gnet track pro, telkomsel and indosat each have a bad reputation in tanjungbungin village, while in solokan village, telkomsel and indosat are bad, respectively. in telukbuyung village, telkomsel and indosat each have a bad reputation. for ping rate performance with kpi jitter standards, telkomsel and indosat each have a bad and sufficient predicate in tanjungbungin village, while in tanjungbungin village,solokan villages telkomsel and indosat each have a bad reputation. in telukbuyung village, telkomsel and indosat have good and bad predicate respectively. 3.4 discussion of voice sequence in the voice sequence, the parameters tested are accessibility which includes call attempts, successful calls, and dropped calls. the result is that with the kpi accessibility standard, telkomsel and indosat's voice call performances each have a ratio of 100% in tanjungbungin village, while in solokan village, telkomsel and indosat have a ratio of 100% and 98%, respectively. for telukbuyung village, telkomsel and indosat have ratios of 100% and 84%, respectively. iv. conclusion the results of the comparison of performance between service providers can be analyzed from the conclusions obtained. indosat tends to be superior based on signal quality even though it gets a bad predicate, but telkomsel has poor results in telukbuyung village. in terms of data access performance, indosat is superior in tanjungbungin village, in the other two villages. telkomsel is slightly superior in terms of upload and download rate. however, for ping quality, indosat is superior in tanjungbungin and solokan while telkomsel is superior in telukbuyung village. for voice call quality, telkomsel is superior with a call success ratio of 100%. meanwhile, indosat only obtained a 100% ratio in tanjungbungin. for cellular operators, which in this case become the object of the author's research, to be able to improve the quality of their services. in this case the signal quality and the placement of the base transceiver station (bts) location, in order to adjust to the natural topology and the surrounding environment. if possible, bts placement can be done in areas with high topology in order to reduce obstacles from surrounding objects and are in areas with high population density. this can be applied to the village which is a weakness in this study, both telkomsel and indosat operators. for operator service users who are the object of this research, namely telkomsel and indosat. in order to be able to make this research as a reference material to determine the operator in accordance with the quality of service in each region. references [1] m. v. panjaitan and a. a. zahra, “analisis quality of service (qos) jaringan 4g dengan metode drive test pada kondisi outdoor menggunakan aplikasi g-nettrack pro,” transient, vol. 7, no. 2, p. 409, 2018. [2] irwansyah, “analisis kualitas koneksi jaringan internet 4g xl dan smartfren di wilayah kota palembang,” pros. semin. nas. pendidik. tek. inform., vol. 8, pp. 73–79, 2017. [3] a. sugiharto and i. alfi, “komparasi performa jaringan antara penyedia layanan seluler 4g lte di area kota yogyakarta,” angkasa, vol. 11, no. 1, 2019. [4] m. prakoso, f. rofii, and a. qustoniah, “aplikasi drive test berbasis android pada jaringan seluler 3g dan 4g,” widya tek., vol. 26, no. 1, pp. 113–128, 2018. [5] i. g. made, y. priyandana, a. saputra, and p. k. sudiarta, “analisis hasil drive test menggunakan software g-net dan nemo di jaringan lte area denpasar,” e-journal spektrum, vol. 5, no. 2, pp. 216–223, 2018. [6] solehudin, arip, bayu priyatna, and nono heryana. "analysis effect of zfone security on video call service in wireless local area network." paper title (use style: paper title) vol. 4, no.2 july 2023 | 63 saubhagya: an online food donation platform for ending hunger and malnutrition in sri lanka g.h.t.r. irushika1, j.j. sathsara2, k.v.m. wijesinghe3, d. i. de silva4, p.k.i. udeshika5, r.r.p. de zoysa6 1,2,3,4,5,6department of information technology sri lanka institute of information technology new kandy rd, malabe, sri lanka 1ruchiniirushika@yahoo.com, 2jithmisathsara098@gmail.com, 3maheshiw99@gmail.com, 4dilshan.i@sliit.lk, 5 ishaniudeshika16@gmail.com, 6 rivoni.d@sliit.lk. abstract hunger and malnutrition continue to be significant challenges in developing nations, including sri lanka. to address this issue, the research paper presents "saubhagya," an online web application that provides a platform for social assistance. the platform allows individuals to donate food and groceries to needy organizations such as blind, deaf, orphanages, making it a user-friendly and effective solution. users are required to register as food donators, needy people(organizations), partners and food collection agents. the system connects these user groups when necessary, ensuring a smooth and efficient process. one unique feature of "saubhagya" is its live capability of tracking food collection and delivery using gps and google maps. this feature ensures that food donations are delivered to the right organizations promptly, promoting transparency, accountability, and communication among users. the research paper aims to evaluate the effectiveness of "saubhagya" in reducing hunger and malnutrition in sri lanka through user feedback and system performance metrics. if successful, the platform can be scaled to other developing nations facing similar challenges. this research demonstrates the potential of digital platforms in addressing social and environmental challenges. by leveraging technology, collective action can be harnessed to create positive social impact. "saubhagya" represents a significant step forward in the fight against hunger and malnutrition, and it is hoped that it can inspire others to use technology to address pressing global issues. keywords : food donation management, hunger and malnutrition, food insecurity, donator, food donation web application abstrak kelaparan dan kekurangan gizi terus menjadi tantangan yang signifikan di negara-negara berkembang, termasuk sri lanka. untuk mengatasi masalah ini, makalah penelitian menghadirkan "saubhagya", sebuah aplikasi web online yang menyediakan platform untuk bantuan sosial. platform ini memungkinkan individu untuk menyumbangkan makanan dan belanjaan ke organisasi yang membutuhkan seperti tunanetra, tuli, panti asuhan, menjadikannya solusi yang ramah pengguna dan efektif. pengguna diharuskan mendaftar sebagai donatur pangan, orang (organisasi) yang membutuhkan, mitra dan agen pengumpul pangan. sistem menghubungkan kelompok pengguna ini bila diperlukan, memastikan proses yang lancar dan efisien. salah satu fitur unik "saubhagya" adalah kemampuannya melacak pengumpulan dan pengiriman makanan secara langsung menggunakan gps dan google maps. fitur ini memastikan bahwa donasi makanan dikirimkan ke organisasi yang tepat dengan segera, mempromosikan transparansi, akuntabilitas, dan komunikasi antar pengguna. makalah penelitian bertujuan untuk mengevaluasi keefektifan "saubhagya" dalam mengurangi kelaparan dan kekurangan gizi di sri lanka melalui umpan balik pengguna dan metrik kinerja sistem. jika berhasil, platform tersebut dapat ditingkatkan ke negara berkembang lainnya yang menghadapi tantangan serupa. penelitian ini menunjukkan potensi platform digital dalam mengatasi tantangan sosial dan lingkungan. dengan memanfaatkan teknologi, tindakan kolektif dapat dimanfaatkan untuk menciptakan dampak sosial yang positif. "saubhagya" p-issn : 2715-2448 | e-issn : 2715-7199 vol.4 no.2 july 2023 buana information technology and computer sciences (bit and cs) mailto:1ruchiniirushika@yahoo.com mailto:jithmisathsara098@gmail.com mailto:maheshiw99@gmail.com mailto:4dilshan.i@sliit.lk mailto:ishaniudeshika16@gmail.com mailto:rivoni.d@sliit.lk vol. 4, no.2 july 2023 | 64 merupakan langkah maju yang signifikan dalam memerangi kelaparan dan kekurangan gizi, dan diharapkan dapat menginspirasi orang lain untuk menggunakan teknologi guna mengatasi masalah global yang mendesak. kata kunci : pengelolaan donasi pangan, kelaparan dan gizi buruk, kerawanan pangan, donor, aplikasi web donasi pangan i. introduction hunger and malnutrition are critical challenges faced by developing nations worldwide. despite significant economic development and increased agricultural output, starvation and malnutrition continue to be major obstacles in many countries, often due to environmental degradation, drought, and loss of biodiversity. sri lanka is one such country facing these issues. to address the problem of hunger and malnutrition in sri lanka, this research paper presents "saubhagya," an online web application that provides a platform for social assistance. the platform allows individuals to donate food and groceries to needy organizations such as blind, deaf, orphanages, and more. the goal of this platform is to address the issue of hunger and malnutrition in sri lanka by providing an efficient and user-friendly solution. the problem statement of this research is to address the issue of hunger and malnutrition in sri lanka. the significance of this research lies in the potential of digital platforms to address social and environmental challenges, particularly in developing nations. the research questions for this study are: how effective is "saubhagya" in reducing hunger and malnutrition in sri lanka?, what are the user feedback and system performance metrics of the "saubhagya" platform?, how can "saubhagya" be scaled to other developing nations facing similar challenges?, the remaining sections of this paper will discuss the design and implementation of "saubhagya," the evaluation of its effectiveness, and its potential impact on addressing hunger and malnutrition in sri lanka and other developing nations. the paper concludes by highlighting the significance of digital platforms in addressing global issues and the potential for future research in this area. ii. litreture review hunger and malnutrition are persistent problems in many developing nations, including sri lanka. a growing body of research indicates that digital platforms can play a crucial role in addressing these challenges by facilitating efficient food donation systems. one study [1] found that digital platforms can help bridge the gap between food waste and food insecurity by providing a direct link between food donors and needy organizations. the study highlighted the importance of technology in reducing food waste and alleviating hunger. similarly, a study [2] suggested that digital platforms can be effective in addressing malnutrition by enabling real-time tracking and monitoring of food distribution. the study found that digital platforms increased the efficiency and transparency of food distribution, leading to improved nutritional outcomes. other studies have highlighted the role of digital platforms in promoting social and environmental sustainability. for instance, a study [3] emphasized the importance of digital platforms in reducing food waste and promoting sustainable consumption practices. overall, the literature suggests that digital platforms have great potential in addressing social and environmental challenges, including hunger and malnutrition. "saubhagya" represents an innovative application of digital technology in addressing the problem of hunger in sri lanka, and it is hoped that this research will contribute to the growing body of knowledge on the effectiveness of digital platforms in social and environmental problem-solving. iii. related work 1. share the meal’ donation and fundraising application 'share the meal' is a contribution and fundraising online application [4] that allows users to give money to purchase a meal for someone in need. there are also fundraisers set up to assist people and to promote sustainable agriculture. users can give money for a single meal, with a minimum donation of $0.80 usd. the problem with this service is that users may only contribute money to others in order to buy a meal; they cannot personally contact the individual and assist them. 2. food donation connection’ food wastage management application 'food donation connection' is a website [5] that handles extra food and distributes it to those in need. 'food donation connection' organizes food donation programs for restaurants interested in donating food. donors vol. 4, no.2 july 2023 | 65 receive financial rewards through tax savings, as well as engagement in community and business reputation, as part of the charity process. the issue with this application is that only restaurants may be donors, and they must also collaborate with a charitable organization. 3. application for reducing food waste in india (no food waste) it is a food collection and waste management network [6] that gathers excess food from individuals and companies and distributes it to those in need. they even gather food from celebrations such as weddings, parties, and other special occasions. the group only needs a phone call with the location, and they will come and pick up the food. they also have a transport vehicle called 'foodiva,' which was designed just for collecting leftovers. 4. application for reducing food waste in india (no food waste) food insecurity affects around 34 million individuals in the united states [7]. 'feeding america' is an organization that connects people with food banks, food pantries, and local food programs. it is beneficial to link with local food banks so that people may find their location and pick up some food. donors can also find out where food banks are located and give food to them. the application is only accessible in the united states. iv. methodology saubhagya: food donation web application" is made up of the following components: react, a frontend javascript toolkit for developing web apps; mongodb, a document-based open-source database; and nodejs is a javascript runtime environment, and express is a nodejs web application framework. figure 1: mern stack development architecture figure 1 depicts the essential components of a full stack application. a full stack application is made up of two major components: the frontend and the backend. the code that shows the application in the user's browser is typically referred to as the frontend. the backend is made up of all the logic that represents the program executing on servers and connecting to the database. the frontend is often built with reactjs. everything related to showing data to the user, re-rendering objects in the browser using dom elements, and furthermore is written by the developer. all http requests from the frontend are handled by the backend server, which also loads data from the database. the database is a location where all of the data from the application is saved in schemas.the frontend of the saubhagya web application is reactjs [8]. react is a popular javascript library for creating user interfaces. it is one of the quickest and most versatile javascript libraries for developing apps, and it has a large community to assist fellow developers if they run into any issues. reactjs allows developers to design applications based on components. developers split the whole user interface into reusable components since programmers break down difficult code into smaller bits. the various components are subsequently merged into parent components, which are then displayed. to study reactjs, programmers must have a fundamental understanding of html [9], css [10], and javascript [11]. html is utilized in the creation of web applications. it is equivalent to the human body's skeleton. css is an abbreviation for cascading style sheet. it is used to create and style websites, and it allows developers to employ colors, fonts, and a variety of other aesthetic options. javascript, a text-based computer language that can be used on both the client and server sides, may be used to create interactive web pages. after being acquainted with these web development languages, programmers must understand npm. npm is a package manager for node.js. it is used to add node modules and packages to the project. fundamental reactjs concepts such as component architecture, state, props functional component, class, and how to integrate api with react app, among others, are required when creating a reactjs-based web application. react router will assist vol. 4, no.2 july 2023 | 66 developers in loading specified interface material and redirecting to certain sites. server render and webpack assist developers in keeping dependencies in a project's static file. loaders in webpack assist to conduct certain tasks in the project. figure 2: nodejs: what is nodejs for figure 2 shows that nodejs [12] is neither a language nor a framework. it is a javascript runtime that is free and open source. the javascript engine, like nodejs, executes code in the runtime environment, although it contains certain extra server-side modules. because nodejs is developed in c, c++, and javascript, it is highly quick and has excellent performance. nodejs has a lot of functionality, which is why it can run javascript code on the server. figure 3: restful api using node.js and express.js expressjs [13] is a nodejs framework built for fast and simply creating apis, online applications, and cross-platform mobile apps. it has excellent performance; it is quick, light, and unopinionated. express does not compel programmers to create code in a specific way. expressjs is a scripting language for the server. if programmers must develop a rest api, expressjs will shorten the time required to code. that is why the expressjs framework for nodejs was created. this is seen in figure 3. mongodb [14] is a document-oriented, no-sequence database (nosql). it was originally made accessible in august 2009. when it comes to relational data structures, mongodb substitutes documents for the rows that are common in such models. because of its adaptability, developers can deal with evolving data models. mongodb supports embedded documents, arrays, and other document-based capabilities and may define complex hierarchical connections with a single record. the document is schema-free since the defined keys are not fixed. vol. 4, no.2 july 2023 | 67 large-scale data transfers are thus out of the question. before beginning with backend implementation, programmers must install the necessary applications. 1. vs code or a similar editor 2. the most recent node.js version 3. api postman creating frontend install node and npm first to rapidly set up a development environment for the application. visit the nodejs website and download version 18.15.0 lts. the frontend of the project may now be created by developers. the command run "create-react-app saubhagya-web-application" may be used for this. this will create a folder and install the packages required to run react in that folder. within that folder, type npm start, and a local server will be launched. creating backend 1. make a new folder and name it, then open it in vs code and run the command "npm init -y" to start the project. 2. in the terminal, type "npm i cors dotenv express mongoose" to install the dependencies. 3. modify the server.js file's main entry point. 4. make the folders database, controllers, models, and routers. figure 4 depicts the previously described folder arrangement. figure 4: folder structure of the backend 5. replace the scripts with the ones displayed in figure 5. figure 5: backend package.json file 6. to begin the server, programmers must import express, then configure the app with express(), create a get function for the endpoint "http://localhost:3000" with app.get(), and set the port to 3001. vol. 4, no.2 july 2023 | 68 7. run the nodemon server with the command "npm run dev." if the server starts properly, the terminal should display "server is running on port 3000." 8. copy the mongodb url from the database and put it into the.env file. replace and with the database's login and password. 9. open the database folder's index.js file and import mongoose, the database url from the .env file, construct the database connection function for connecting to the database, export the method, and call it in the server.js file. 10. in the models folder, create different files that specify the database schemas for the services we provide in this application, such as boarding schema, product schema, and etc. 11. define the end point methods, such as create, update, delete, and retrieve methods, and etc. 12. include the route end points in the server.js file. 13. use the postman api to test the newly generated end points [15]. tools used for the implementation azure board [16] is one of the key azure devops services that is used in a project to track work with kanban boards, making it simpler to finish tasks as a team. when programmers utilize azure devops, they will have an azure board on which they may construct and configure a kanban board. a kanban board contains information on all of the work items that developers will deliver or work on in a certain project. it is simple to manage and work on items in the project backlog using the azure board. the backlog will contain all of the tasks that must be completed in the current or forthcoming sprint. azure boards will also give a variety of reports. as the project progresses, developers must update their tasks and features on the kanban board. azure boards are great for task management and tracking because they give a clear image of work completed/being completed by team members. github [17] is a web-based graphical user interface hosting service. github enables team members to communicate, track, and update their work while working on the project from any place. as a result, projects remain open and on schedule. the team can stay organized and on the same page by using github. pull requests on github help teams evaluate, improve, and suggest new code. implementations and recommendations can be addressed before making modifications to the source code. sonarqube [18] simply scans through a developer's code and discovers errors early on. it is an open-source static testing analysis application. it is used by developers to maintain the uniformity and quality of their source code. code quality checks look for potential defects, design inefficiencies caused by code errors, duplication of code, insufficient test coverage, and other concerns. selenium [19] is a free (open source) automated testing framework for web applications that runs across several browsers and platforms. selenium's major focus is on web-based application automation. selenium is a collection of tools that each target a different aspect of an enterprise's testing needs. v. proposed system the proposed system "saubhagya” is an online web-based system that serves as a platform for people who want to donate food to those in-need with the goal of zero hunger in sri lanka. this proposed system will eliminate all disadvantages of the current food donation systems and benefit all users. the basic prerequisite to use this web application is a smartphone. the system has four types of users and this system has a unique software engineering feature. the users are, 1. needy people 2. donators 3. partners 4. delivery agents the unique feature of this system is that, the live capability of tracking food collection and delivery using gps and google maps. vol. 4, no.2 july 2023 | 69 figure 6: home page-saubhagya all the four types of users should login or register with the system to utilize saubhagya. needy people needy people is a key user of the system. they can present their organization in the system to obtain food donation via the system. needy people organizations are viewed many donators who visit the web application. initially, needy people must register their organization in the system. they should provide details such as organization name, address, contact number, email, number of children, number of adults, meals, food preferences, other required necessities and organizational logo. figure 7: use case diagram of needy people the above use case [figure 7] shows the needy people functional activities. these are the basic functionalities that needy people can operate in the web-based system. needy people can post requests of food to obtain it from the donators. also, the donators can send food donation requests to the needy people. there needy people can accept or decline the food donate requests through the system. vol. 4, no.2 july 2023 | 70 figure 8: needy people make food request needy people can do all the curd operations through the system [figure 8]. such as insert the organization to the system. update the organization details, view or delete organization from the system. search other needy people organizations. generate reports on food requests. when considering the advantages of the system for needy people, saubhagya is a life savior for the needy people because physical food donation programs are not happening as expected. to get rid of that problem, this online food donation platform for the needy people is the ideal option to eliminate hunger in sri lanka. because, multiple food donators can connect with this platform. donators the key resource of this system is the donors. they give food and other necessities to those in need. donators might donate for the food requests raised by the needy people. additionally, they can use the system to contribute food to their preferred organization for those in need. the main benefit of this process is that it enables donors to connect with partners to post about their spare food and then donate it to those in need who cannot afford to buy food for themselves. unlike other platforms for food donations, saubhagya gives donors the opportunity to make donations and manage food requests. figure 9: use case diagram of donator the functional actions of the donor are shown in the above-mentioned use case [figure 9]. these are some of the fundamental features of the web application that donors can use. in addition to the benefits of the system for vol. 4, no.2 july 2023 | 71 donations, it also saves donors time because physical donation programs can occasionally be a little timeconsuming. the ideal approach to manage donations is through the web platform, which will eliminate that issue. the user satisfaction is yet another benefit in addition to those mentioned before. figure 10: donator dashboard this interface [figure 10] displays the donators' dashboard. the user can view the active donations, pending donations, and reject donations. users can also view the latest requests received for his/her donations. donators can do all the curd operations through the system. such as insert the donations to the system. update the donation details, view or delete donations from the system. generate reports on food donations. partners (food donor partner) a food donor partner is a person who voluntarily supports donations. they can support a donation in a variety of ways. those are: • by donating goods etc. • donation of processed food. • by donating money. • by offering discounts on the purchase of goods or food (only for business locations). figure 11: key user flow of donor partner vol. 4, no.2 july 2023 | 72 before using the web application, the user must register and log in (if a new user). after login, the user will be redirected to a separate dashboard [figure 11]. there the donor partner can see "needy people” as well as "donations”. under donations, partner can request to contribute to a donation of their choice. figure 12: use case diagram of donor partner above diagram [figure 12] demonstrate the basic functionalities that a partner can do. partners can view donation partnership requests, generate reports, and view the user profile. under “my partnerships”, all data about solicitations and donation partnerships made so far is displayed. he or she can edit a request that has not been approved by the donor. the ability to delete is also present but only requests rejected by the donor can be deleted. and like other features, the donor partner can view the profile, edit the profile, generate reports and delete his or her account. delivery agents figure 13: add food collection details vol. 4, no.2 july 2023 | 73 the proposed system for food collection agents consists of several key features, starting with registration where users can sign up to become food collection agents by providing their name, contact details, and other necessary information. after registration, users can log in to their personalized dashboard, which includes the view my account page and food collecting details. the view my account page allows users to view and update their profile information, while the add new food collecting details page enables users to add new food collection activities to the system [figure 13]. the view food collecting details page displays existing food collection activities, and users can edit or delete them on the update & delete food collecting details page. additionally, the system provides a search function that enables users to search for specific food collection activities based on different parameters, and users can generate reports based on the collected food data. overall, this key user flow diagram outlines how users can navigate through the system to manage their food collection activities efficiently. figure 14: key user flow of delivery agent the above diagram [figure 14] highlights the functionality offered to food collecting agents in order to efficiently handle the food collection process. agents can create new food collecting activities, view current ones, update or delete them, and search for specific activities based on certain criteria. in addition, the system allows the user to generate reports based on the food data you collect. unique feature of saubhagya web application the real-time tracking feature in saubhagya uses gps and google maps to track the food collection and delivery process. when a donator or partner posts a donation or support request on the platform, the system automatically assigns a delivery agent to collect and deliver the donation to the needy people organization. the delivery agent then uses the saubhagya app on their smartphone to navigate to the collection and delivery locations. vol. 4, no.2 july 2023 | 74 as the delivery agent moves from the collection location to the delivery location, the system tracks their location using gps and displays it on a map in real-time. this allows all the four types of users to track the progress of the delivery process. the needy people organization can see the status of the delivery in real-time and estimate the arrival time. the donator or partner can also see the status of their donation or support and track its progress. additionally, the delivery agent can use the real-time tracking feature to navigate to the delivery location more efficiently and update the status of the delivery in the system. the real-time tracking feature is a significant advantage of saubhagya, as it provides transparency and accountability to the food donation process. it allows all the four types of users to track the delivery process in real-time and ensures that the donations reach the needy people organization on time. it also helps to build trust between the users of the platform and encourages more people to donate food and support. overall, the combination of gps and google maps in saubhagya provides an accurate, efficient, and userfriendly way to track the location of delivery agents in real-time. this benefits all the four types of users by providing transparency, accountability, and trust in the food donation process. vi. disscussion the saubhagya web application is a cutting-edge solution to the ongoing challenges of hunger and malnutrition in sri lanka. by leveraging the latest technology and utilizing the mern stack, the application provides a user-friendly and effective platform for individuals to donate food and groceries to those in need. the front end of the application is built using reactjs, which is a popular javascript library that allows for efficient rendering of web pages. the back end of the application is built using expressjs and node.js, which are two powerful and widely-used web development frameworks. these technologies enable the application to handle large amounts of data and perform complex calculations in real-time. to store and manage the data, the application uses a mongodb database. this database is highly scalable and can handle large amounts of data, making it an ideal choice for a platform like saubhagya. the application also uses various libraries such as axios, mongoose, and express, which help to streamline the development process and improve the performance of the application. in terms of design, the saubhagya web application has been developed with the user in mind. the interfaces are simple and intuitive, featuring side navigations, menus, icons, and buttons that make it easy for users to navigate and understand the application. additionally, the application utilizes external css libraries such as bootstrap and material icons to make it more visually appealing and engaging for users. overall, the saubhagya web application is a powerful and innovative solution to the challenges of hunger and malnutrition in sri lanka. by utilizing the latest technology and developing a user-friendly platform, the application has the potential to make a significant impact in reducing hunger and improving the lives of those in need. vii. conclusion in conclusion, the issue of hunger and malnutrition remains a significant challenge in developing nations such as sri lanka. the research paper proposes "saubhagya," an online web application that provides a platform for social assistance to address this issue. this system allows individuals to donate food and groceries to needy organizations such as blind, deaf, orphanages, and more, making it a user-friendly and effective solution. the unique feature of "saubhagya" is its live capability of tracking food collection and delivery using gps and google maps, ensuring transparency, accountability, and communication among users. the paper aims to evaluate the effectiveness of "saubhagya" in reducing hunger and malnutrition in sri lanka through user feedback and system performance metrics. the significance of this research lies in the potential of digital platforms to address social and environmental challenges, particularly in developing nations. by leveraging technology, collective action can be harnessed to create positive social impact. the findings of this research will contribute to the development of digital platforms that can address pressing global issues. the proposed system "saubhagya" has the potential to make a significant impact on reducing hunger and malnutrition in sri lanka and other developing nations. the system's scalability makes it a promising solution that can be adapted to other countries facing similar challenges. this research demonstrates the potential of technology to address social and environmental challenges and inspire further research in this area. viii. references [1] kusuma, s. s., & dileepkumar, k. r. (2020). food donation and waste management system: a survey. in 2020 international conference on smart electronics and communication (icosec) (pp. 1-5). ieee. vol. 4, no.2 july 2023 | 75 [2] nair, p. r., & george, b. (2018). socio-economic impact of tourism: a case study of kerala. indian journal of tourism and hospitality research, 5(1), 27-37. [3] oktaviani, r., & salim, f. d. (2020). digital platform for donating food: design and development of foodish. in 2020 6th international conference on science in information technology (icsitech) (pp. 300305). ieee. [4] share the meal. available at: https://sharethemeal.org/. accessed on: oct. 2023. [5] food donation connection. available at: https://www.foodtodonate.com/. accessed on: oct. 2023. [6] https://nofoodwaste.org/. accessed on: oct. 2023. [7] feeding america. available at: https://www.feedingamerica.org/. accessed on: oct. 2023. [8] react. (n.d.). react a javascript library for building user interfaces. retrieved from https://reactjs.org/. [9] w3schools. (n.d.). html tutorial. retrieved from https://www.w3schools.com/html/. [10] w3schools. (n.d.). css tutorial. retrieved from https://www.w3schools.com/css/. [11] w3schools. (n.d.). javascript tutorial. retrieved from https://www.w3schools.com/js/. [12] node.js. (n.d.). node.js. retrieved from https://nodejs.org/en/. [13] express. (n.d.). express.js. retrieved from https://expressjs.com/. [14] mongodb. (n.d.). mongodb the most popular database for modern apps. retrieved from https://www.mongodb.com/. [15] postman. (2021). postman | the collaboration platform for api development. retrieved from https://www.postman.com/. [16] azure devops. (n.d.). boards azure devops. retrieved from https://azure.microsoft.com/enus/products/devops/boards/. https://www.w3schools.com/css/ https://www.w3schools.com/js/ vol. 4, no.2 july 2023 | 32 j.valarmathi1, v.t.kruthika2 a philosophical study of agricultural image processing techniques j.valarmathi1, v.t.kruthika2 1,2assistant professor, pg and research department of computer science and applications, vivekanandha college of arts and sciences for women(autonomous), tiruchengode – 637 205, india. 1valarmathij@vicas.org, 2kiruthika@vicas.org abstract the development of agriculture in china has been substantially aided by the development of image processing technologies. it is simple for people to comprehend the significance of image processing technology for agricultural development by presenting the application status of image processing technology in agriculture and its impact on agricultural production value. this research examines how image processing technology is used in agriculture on the basis of that information. this study first examines how image processing technology is used in the world of agriculture. second, this study applies both classic machine recognition technology and image processing technology to crop pest identification, analyses their effects, and highlights the application effect of image processing technology in the agricultural industry. the findings indicate that this approach has a recognition rate of 86%, 89%, 91%, 83%, 78%, and 79%, respectively. it is evident that the detection of crop diseases and insect pests is improved by the use of image processing technology. keywords: agriculture, diseases, insect pests, image processing, machine recognition abstrak perkembangan pertanian di china telah banyak dibantu oleh perkembangan teknologi pemrosesan citra. sangat mudah bagi orang untuk memahami pentingnya teknologi pengolahan citra untuk pembangunan pertanian dengan menyajikan status penerapan teknologi pengolahan citra di bidang pertanian dan dampaknya terhadap nilai produksi pertanian. penelitian ini mengkaji bagaimana teknologi pemrosesan citra digunakan dalam pertanian berdasarkan informasi tersebut. penelitian ini pertama kali mengkaji bagaimana teknologi image processing digunakan dalam dunia pertanian. kedua, penelitian ini menerapkan teknologi pengenalan mesin klasik dan teknologi pemrosesan gambar untuk identifikasi hama tanaman, menganalisis pengaruhnya, dan menyoroti efek penerapan teknologi pemrosesan gambar dalam industri pertanian. temuan menunjukkan bahwa pendekatan ini memiliki tingkat pengenalan masing-masing 86%, 89%, 91%, 83%, 78%, dan 79%. jelas bahwa deteksi penyakit tanaman dan serangga hama ditingkatkan dengan penggunaan teknologi pemrosesan gambar. kata kunci: pertanian, penyakit, serangga hama, image processing, pengenalan mesin p-issn : 2715-2448 | e-issn : 2715-7199 vol.4 no.2 july 2023 buana information technology and computer sciences (bit and cs) mailto:1valarmathij@vicas.org mailto:kiruthika@vicas.org vol. 4, no.2 july 2023 | 33 i. introduction the first use of digital image processing was in 1920. the time it takes for a photograph to be transferred from the atlantic side has decreased from more than one week to three hours since the implementation of batland's cable image transmission system [1-2]. image processing technology has advanced with the quick development of computers [3–4]. humans rely heavily on images to exchange and collect information. every element of human work and life has steadily been impacted by the application of picture analysis and processing. the development of this technology is developing by leaps and bounds, and its application domains are continuously expanding along with the breadth of human activities expanding and the ceaseless appearance of scientific theories, and its great achievements are also accumulating day by day [5-6]. the efficiency of physical labour was historically quite poor, and agricultural growth trends were difficult to manage [7]. according to their work experience, the staff discovered that monitoring the growth of crops can provide them with information about the soil, water, and air humidity, which they can use to improve subsequent work appropriately and ensure that the crops grow normally and produce higher overall output values [8]. as science has advanced quickly, digital image processing technology has become increasingly widely used in agriculture. examples include the detection of pesticides in vegetables, the control of pests in crops, the identification of crop growth trends, and the colour coding of crops. digital image processing technology plays an indispensable role [9-10]. for the development of the agricultural field, it is therefore crucial to examine the image processing technology's current state of use. this research paper provides a quick overview of image processing technologies. four categories of image processing technology methods are covered in this article: image denoising, image rectification, image segmentation, and picture feature extraction. the implementation of image processing technologies in the sphere of agriculture is also examined in this research. additionally, the impacts of applying standard machine recognition technology and image processing technology to crop pest identification are compared in order to emphasise the application effect of image processing technology in the agricultural industry. the findings of the trial demonstrate how much more effective the image processing-based detection method is. ii. overview of image processing technology image denoising, correction, segmentation, feature extraction, and other techniques are all examples of image processing technology. a. image denoising noise is introduced and the image signal is contaminated during the processes of image acquisition and transmission due to the influence of equipment and external variables. denoising is a well-known and fundamental issue in image processing and analysis. gaussian white noise and salt-and-pepper noise are two types of typical picture noise. b. image correction image skew correction is the process of recovering the image that does not comply with the standard before processing the image. image skew is primarily caused by the deviation of scanning layout during the process of acquisition. it is highly typical to have diverse image effects during the image acquisition process, particularly the image tilt, due to the varied technology used by each individual. if there is no pertinent tilt correction applied before processing an image with skew, the final automatic recognition will be severely hampered. the image is first examined, and the degree of image skew is determined. then, to perform the picture rectification, the same angle is rotated in the original coordinate in vol. 4, no.2 july 2023 | 34 accordance with the determined tilt angle. the projection method, nearest neighbour method, hough transform, radon transform, and more methods are currently available. c. image segmentation one of the challenges in image processing and a significant issue in picture analysis is image segmentation. nevertheless, picture segmentation is the initial stage of the image processing process. the retention and display of image features after segmentation have a significant impact on the subsequent image processing. image segmentation can be considered to have a direct impact on the outcomes of image analysis and processing, and effective image segmentation will set up the final image processing on a solid foundation. the most well-known and often used segmentation technique is threshold-based segmentation. d. image feature extraction the pixel (x, y) in the image is assumed to be integrated throughout the image feature extraction process, and the total of all pixels above the pixel is represented by a sub integration system. iii. methodology this research applies image processing technology and classical machine recognition technology to crop pest identification, respectively, and evaluates their results in order to highlight the application effect of image processing technology in the agricultural industry. a. subjects two crops that were the same size were chosen as the experimental objects in this paper. the two fields' agricultural development patterns were comparable, they were in the same area, and other circumstances were largely the same. b. test object this investigation's goal is to find crop pests. pests and illnesses, such as the alfalfa armyworm, blue grey butterfly, bean grey butterfly, flame noctuid, and bean leaf roller, were chosen as the detection items in this study. c. detection index the detection index in this research is based on the degree of recognition for the two approaches. the detection effect improves with higher recognition degrees. iv. results and discussions a. analysis of image processing technology in agricultural field in this paper, the application of image processing technology in agricultural field is analyzed, agriculture uses image processing technology extensively. the prevalent uses of image processing technology in agriculture are examined in this research. according to table 1 and figure 1, the current usage of image processing technology in agriculture focuses mostly on crop growth monitoring, disease vol. 4, no.2 july 2023 | 35 and insect pest diagnosis, nutritional status monitoring, crop maturity monitoring, and crop colour identification. in order to effectively assess the growth state of crops, 29.3% of them are utilised to monitor crop growth. figure 1. aplication of image processing technology in agriculture 14.5% of cases of illnesses, insect pests, and weeds are diagnosed. it is used to provide crops with additional nourishment and water as needed. 16.7% of applications for crop maturity monitoring are made with the goal of increasing crop production effectiveness. crop colour identification employed 17.4% of the crop's colour, and its application goal was to categorise crops. this study then goes into further detail about this. b. monitoring crop growth in general, computer vision technology can be utilised extensively during the entire process of plant growth, monitoring plant growth and development, and if abnormal conditions are identified, it is useful to remedy the problem as soon as possible. crop leaf thickness, rhizome length, and water content are the key monitoring targets, and all pertinent information is meticulously documented. when paired with final data, we can assess crop production comprehensively; when combined with crop fruit photographs, we can determine at any moment whether the fruits are mature, lacking in food and water. c. diagnosis of diseases, insect pests and weeds in addition to giving crops the nutrients they require to grow properly on schedule, it is also essential to deal with the diseases, pests, and weeds that impede crop growth. in the past, this component of work was greatly influenced by agriculture's poor production value. with the development of image processing technology, its use in agricultural work has continued. this liberates the laborious statistical work of crops and significantly reduces the difficulty of staff job. in order for the personnel to carry out preventive work, image processing can be used to forecast potential issues that could arise during the early stages of crop growth. d. monitoring nutritional status real-time photographs of crop leaves and rhizomes can be captured using image processing technology during the growing process, allowing for the monitoring of crop leaf size and rhizome thickness. through the monitoring data, crop-related data can be compared to the average state to determine whether there are any nutritional deficiencies or other problems. this allows for the timely vol. 4, no.2 july 2023 | 36 development of an effective remediation plan, which ensures that crops grow normally and receive enough water and nutrition. e. monitoring maturity the crop fruit picture points may be gathered from a wide range using browser image and other relevant analysis technologies, and the crop growth and maturity can be precisely determined by the obtained parameters. these technologies allow us to assess the fruit's maturity and create efficient defences. for instance, the earlier-maturing fruit can be plucked earlier to prevent decay and other decline, which is helpful for the systematic management of the fruit situation and increases the effectiveness of production. f. identify crop colors the visual characteristic of colour makes it simple to assess the quality of crops. digital imaging technology's gathering and analysis of colour traits transform into a detection method to determine whether the crops are of excellent quality. a theoretical basis was established for the systematic and uniform development of maize quality inspection by using the detection of corn quality as an example. the analysis of multiple image indicators, such as colour saturation and sensitivity of corn kernels, can be used as the quality grading standard of corn kernel sweetness and fineness. g. analysis of image processing technology in agricultural field the traditional machine recognition technology and the image processing technology detection method are used to identify crop diseases and insect pests in order to study the application effect of image processing technology in agricultural fields, and the application effect of the two methods is compared. in this study, six different illnesses and insect pests—the bean leaf roller, bean leaf borer, flame armyworm, blue grey butterfly, bean grey butterfly, and alfalfa armyworm—were chosen as detecting objects. table 2 and figure 2 show that there are some discrepancies in the recognition rates of the two distinct detection methods for crop diseases and insect pests. vol. 4, no.2 july 2023 | 37 the classic machine recognition technique among them has a recognition rate of 65%, 71%, 74%, 63%, 64%, and 62%, respectively. additionally, this approach had a recognition rate of 86%, 89%, 91%, 83%, 78%, and 79%, respectively. the detection approach utilising image processing technology can more successfully detect illnesses and pests based on the recognition data of the two detection methods. v. conclusions in order to achieve the modernization of agriculture level, digital image processing technology is widely applied in all facets of agriculture. despite a late start, the use of image processing technologies in chinese agriculture still produced positive outcomes. in this essay, the use of image processing technology in agriculture was examined, and its effects were researched. this study demonstrates how image processing technology is mostly employed in agriculture for the following five purposes: crop growth monitoring, disease and insect pest detection, maturity monitoring, and crop colour identification. additionally, the use of image processing technology in agriculture has produced positive outcomes, which encourages the growth of agriculture. references [1] anup vibhute, bodhe, "applications of image processing in agriculture: a survey", international journal of computer applications, volume 52, no.2, august 2012, pp:34-40. [2] latha, poojith, amarnath reddy, vittal kumar, "image processing in agriculture", international journal of innovative research in electrical, electronics, instrumentation and control engineering, vol. 2, issue 6, june 2014, pp:1562-1565. [3] nan xu, "image processing technology in agriculture", the 2nd international conference on computing and data science (conf-cds 2021), pp:1-7. [4] kousalya,kabilesh,mohan prasath,jayapriya, "image processing techniques in agriculture for plant disease detection and weed detection", international journal of creative research thoughts (ijcrt), volume 9, issue 3, march 2021,pp: 4643-4650. [5] sanjay b. dhaygude, nitin p.kumbhar, "agricultural plant leaf disease detection using image processing", international journal of advanced research in electrical, electronics and instrumentation engineering, vol. 2, issue 1, january 2013, pp:599-602. vol. 4, no.2 july 2023 | 38 [6] prakash, saravanamoorthi, sathishkumar, parimala, "a study of image processing in agriculture", int. j. advanced networking and applications,volume 09, issue 01,2017, pp:3311-3315. [7] riya desai, kruti desai,shaishavi desai, zinal solanki, densi patel, vikas patel, "removal of weeds using image processing: a technical review", international journal of advanced computer technology (ijact), pp:27-31. [8] batmavady, samundeeswari, "detection of cotton leaf diseases using image processing", international journal of recent technology and engineering (ijrte), volume 8, issue 2s4, july 2019, pp:169-173. p-issn : 2715-2448 | e-isssn : 2715-7199 vol.3 no.2 july 2022 buana information tchnology and computer sciences (bit and cs) 41 | vol.3 no.2, july 2022 systematic literature review implementation of the internet of things (iot) in smart city development gunawan 1 study program technical information stmik ymi tegal & stkip invada cirebon email: gunawan.gayo@gmail.com wilda shabrina 2 study program technical information stmik ymi tegal email: wildashabrina8@gmail.com ‹β› wresti andriani 3 study program technical information stmik ymi tegal email: wresty.andriani@gmail.com abstract. internet of things (iot) "the smart city concept has become a dream that big cities in indonesia want to achieve, basically the smart city concept focuses on developing the human element using technology. iot opens up many opportunities for new services by connecting the physical world and the virtual world to various electronic devices. iot aims to leverage cuttingedge technology to support sustainable services for governments and their citizens. systematic literature review (slr) is one method in conducting an overview of previous interrelated researchers. the purpose of the literature study in this research is to understand trending research topics, methods, and architecture in the development of smart cities with iot. in this study, various interesting information was found in various research journals regarding the role of iot in building a smart city that has a role to serve smart city infrastructure, identify and analyze trends, develop smart cities and become innovations for iot. based on the research, there are things that need to be done to improve the study to focus more on the role of iot that can be more utilized. iot is also one of the technologies that are widely used in several countries in the development of iot in the future. iot in smart cities will become an inseparable technology because humans will increasingly depend on iot in their daily lives. keywords: internet of things, iot, smart city, systematic literature review i. introduction the smart city concept has become a dream that big cities in indonesia want to achieve [1]. smart city-based development has become a trend in cities or regional development around the world, with the belief that regions or cities and districts throughout indonesia need to adapt, smart city emphasizes the importance of innovation in solving problems in every city by using information and communication technologies. (ict), sensors, and data analysis as supporting factors and facilitating problem solving (realization factor [2]. basically, the smart city concept focuses on developing the human element using technology. the internet of things (iot) opens up many opportunities for new services by connecting the physical world and cyberspace to various electronic devices in homes, cars, roads, buildings and many other places is recognized as a major research and innovation stream. a decentralized environment, urban iot aims to leverage cutting-edge technology to support sustainable services for governments and their citizens [3]. iot can also be used to increase transparency and promote local government action against residents, increasing public awareness of their condition. therefore, the application of the iot model for smart cities is very attractive to local governments who may be early adopters of the technology, thereby acting as a catalyst for the adoption of the iot model on a larger scale [4]. with all the conveniences that the smart city concept gets by using the internet of things (iot) which already covers various circles, if the current technological sophistication looks like competing in realizing its development. do not escape with all the advantages of the concept with all its sophistication. with all the sophistication of the role of humans as workers can be replaced with machines that have been created. and the worst impact is the loss of jobs and livelihoods for some people, for that we need to be smart enough to be able to go with the flow and take advantage of all the sophistication that exists to make alternative choices for business and develop capabilities [1]. to realize a sustainable city and support improving the quality of life of citizens, but there are shortcomings when implementing iot such as the risk of system errors that may also arise due to power outages, the system also faces cyberattack vulnerabilities [1]. technological developments are now growing rapidly and iot media as a companion will always be present in all devices that are connected to each other, with iot expected to solve these problems [1]. this study aims to see the role of iot in smart cities. thus, the research question is: "what is the role of iot in building a smart city?" ii. method 2.1. study of literature systematic literature review (slr) is one method in conducting an overview of previous studies that are interrelated. slr is one of the standard methods so that the next research process can be re-done by other people with almost the same results. slr is a secondary research to map, identify, evaluate, consolidate and collect information from the main study results related to the research topic [5] 42 | vol.3 no.2, july 2022 the purpose of the literature study in this study is to understand trending research topics, methods, and architecture in predictive analysis with iot. [6] 2.2. research question the purpose of the research question is to maintain the focus of the literature review. this condition facilitates the process of finding the required data. table 1 shows the research questions for this study [6]. table 1 research question id research question motivation rq1 which journals most frequently publish topics on the application of the internet of things (iot) in smart city development? identify the journals that most frequently publish topics on the application of the internet of things (iot) in smart city development. rq2 who are the most active and influential researchers on the topic of implementing the internet of things (iot) in smart city development? identifying the most active and biggest contributors to the topic of the application of the internet of things (iot) in smart city development. rq3 what is the most frequently used topic on application of internet of things (iot) in smart city development? identify the most frequently used datasets in the topic of implementing the internet of things (iot) in smart city development? rq4 what methods are often used for the application of the internet of things (iot) in smart city development? identify methods that are often used for implementing the internet of things (iot) in smart city development. 2.3. search strategy and selection there are two criteria in journal selection, namely inclusion criteria and exclusion criteria. inclusion criteria follow the following points: 1. “smart city”, “internet of thing”, “iot” are in the title. 2. language: english, indonesian. 3. year: 2017 to 2022 4. publication type: journal and book 5. accessibility: documents available in google scholar. 6. document types: pdf, html. exclusion criteria are all journals that are not accessible, all downloaded documents whose publication type does not match the inclusion criteria, all journals with incomplete content, and all journals whose content does not match the theme of the research question. the search process according to stage 4 in the systematic literature review stage above consists of several processes, including the selection of digital libraries and setting keywords. before starting the search, it is necessary to determine or select the appropriate database to find relevant journals. the following are the digital libraries in this study: a. sciencedirect (http; //www.sciencedirect.com/) b. ieee (http://ieeexplore.ieee.org/) c. google scholar (http://scholar.google.com/) d. keywords are developed according to the following steps: 1. identify search terms from the picoc, particularly from populations and interventions; 2. identification of search terms from research questions; 3. identify search terms in the title, abstract and relevant keywords; 4. identification of synonyms, alternative spellings and anonymous of search terms; 5. determination of thorough keywords using the identification of search terms bolean and and or. the keywords used for the search are: 2.4. study selection the main study search and selection process at each stage is shown in figure 1. the selection shown in step 5 was carried out in two steps: exclusion of the main study based on the title and abstract and exclusion of the main study based on the full text. the study selection used was only journals, while books and proceedings were not used in the study selection: 1. english speaking is preferred. 2. journals are included in computer science. 3. we get about 1,000 journals about iot. a. then the selection was made based on the title and abstract of 100 articles. b. the final results of the selection are 15 journals in the main study. 43 | vol.3 no.2, july 2022 figure 1. main study search and selection 2.5. literature extraction and analysis main study ecstasy data extraction was designed to collect data from the main study needed to answer the research question. table 2 data extraction for questions property research question the year of publication and the most active researchers rq1, rq2 datasets used rq3 methods that are often used for implementing internet of things in smart city development rq4 literature analysis a. deep learning analysis originated from the study of artificial neural networks (ann), b. the most popular deep learning methods, especially in smart city research: rnn, cnn, dbn and sae c. there are 4 types of rnn: long short-term memory (lstm), gated recurrent unit (gru), bi-directional rnn and rnn encoder-decoder. d. cnn: consists of a number of convolutional layers connected as hidden layers e. dbn: a generative graphics model that studies the representation of the given data. f. sae: a stacked autoencoder is created by stacking each hidden layer of an ae on top of another. g. rnn and cnn are methods of deep learning that are often used in health research iii. results and discussion from the search results of scientific publications on a list of digital databases that have been selected, several scientific publications have been selected that can be used as references for authors to answer research questions that have been compiled. there are scientific publications that cannot be used as references because they do not specifically discuss the internet of things in smart city development. a summary of the evaluation results is presented in table 3. exclude by full text (20) take the initial list of main studies 17,195 start select digital library list define search character sciencedirect: 1.718 ieee: 99 google scholar: 17.500 exclude by title and abstract (113) make a final list according to major study (15) finish main study found? 44 | vol.3 no.2, july 2022 table 3 summary of evaluation results reference research purposes research result the role of iot [3] showcasing the role iot can play in smart cities. the application component contains the rules governing the decisions made on the smart home control. it is estimated that the application will also receive information from the electric utility regarding the supply of electricity, and information from the weather bureau. play a role in serving smart city infrastructure [4] discusses urban iot design models by describing the specific characteristics of urban iot, and services that can drive the adoption of urban iot by local governments. most smart city services are based on a centralized architecture, where dense and heterogeneous devices placed in urban areas generate various types of data which are then sent via appropriate communication technologies to a control center, where data storage and processing takes place. utilization of iot to make design and other application characteristics more attractive. [6] predict data based on patterns extracted from historical data derived from iot data. there are six research areas that are trending in predictive analytics with iot, namely transportation, agriculture, health, industry. smart home, and environment. identify and analyze trends, methods and architectures used in predictive analytics with iot. [1] providing solutions that make it easier for humans in several sectors of work. infrastructure, management, regulation and citizens are the main components that are very efficient when viewed from the needs of the city of jakarta itself in implementing a smart city. playing a role in the development of a smart city, with a fast internet network, supervision of all cities can be done quickly. [5] applications improve services and help simplify decision making iot with a very heterogeneous network application with various types of objects. provide opportunities and challenges in the decision-making process. the main trends in the application of iot to support decision making are in the health, manufacturing and industry sectors, transportation, agriculture and smart home. [2] creating integration, synchronization, and synergy between smart city development planning at the central and regional levels, encouraging an effective, efficient, inclusive and participatory smart city development process there are 5 factors that support smart cities in temanggung regency, namely nature factors, structural factors, infrastructure factors, superstructure factors, cultural factors. temanggung regency has quite good regional readiness, especially in terms of regional digital infrastructure [7][8] a communications infrastructure that provides unified, simple and economical access to a large number of public services, thereby unleashing potential synergies and increasing transparency to citizens build a multipurpose iot mode even from very limited devices. using information schema & xml schema analyze solutions for urban iot deployments. starting from field trials which are expected to eliminate the uncertainty that still hinders the iot paradigm 45 | vol.3 no.2, july 2022 [9] making the internet more immersive and pervasive. allows easy access and interaction with a wide variety of devices. padova smart city was successfully realized in the city of padova which specializes in the development of innovative iot solutions, which has developed iot mode and control software. analyze available solutions for urban iot implementation. [10] provide an evaluation of the readiness level for implementing smart mobility in jakarta measurement of jakarta's smart mobility is carried out using the indicators contained in smart mobility in jakarta, it can be said that they are ready to implement smart mobility in terms of accessibility and connectivity as well as the use of information and communication technology [11] iot extends the benefits of continuously connected internet connectivity cities in indonesia accept the concept of adaptation from a country that has successfully implemented the concept of a smart city the development of iotbased infrastructure will facilitate the development of smart cities in indonesia. [12] the realization of a smart city to create a city that is comfortable, safe, efficient and sustainable, as well as increasing transparency and efficiency of government governance in the city of bandung the iot concept in bandung is designed based on the results of a review of the theory and the development of previous research iot requires a framework or concept to be implemented because of the many concepts, models and hardware and software devices that have a very broad scope in iot technology. [13] creating an intelligent environment by utilizing various smart objects/devices that have sensory and communication capabilities to generate data and send it over the internet to make decisions an iot system can be represented and described in 3 main layers, namely the transportation layer and the application layer. each has a different technology researchers have shown issues regarding privacy, security and energy management that are still the main focus in the development of iot in the future. [14] iot using fog computing architecture offers a very feasible design study which is expected to be more satisfying with the new concept the results of the smart city design will be divided into 2 parts, namely smart city design in fields and architectural designs in all fields iot is chosen by using fog computing architecture to reduce latency, simplify communication between central devices [15] significant developments in smart cities, iot, crowdsourcing with a focus on technology, challenges and the latest developments related to smart cities. iot development and crowdsourcing are important drivers for smart city development. smart cities act as innovations for iot and crowdsourcing. [16] to synthesize developments in the field of iot and smart cities and advance the existing knowledge base by providing some recommendations for future research. assist researchers who are actively investigating iot in a smart city context. fills gaps in the literature by uncovering various iot and smart city research issues and topics. 46 | vol.3 no.2, july 2022 iv. conclusion in this study, based on the results of slr searches conducted, found various interesting information in various journals that were studied about the role of iot in building smart cities, namely serving smart city infrastructure, identifying and analyzing trends, developing smart cities, and being an innovation for iot. references [1] m. subani, i. ramadhan, a. syah putra, and a. al muslim, “perkembangan internet of think (iot) dan instalasi komputer terhadap perkembangan kota pintar di ibukota dki jakarta,” ikra-ith inform. j. komput. dan inform., vol. 5, no. 1, pp. 88–93, 2021. [2] a. fahrina, y. wirani, a. gandhi, y. ruldeviyani, and y. g. sucahyo, “analisis kesiapan pembangunan smart city daerah studi kasus : kabupaten temanggung,” vol. 9, no. 2, pp. 984–995, 2022. [3] agung ahmad, “pengembangan internet of things pada smart city,” j. sist. cerdas, vol. 1, no. 1, pp. 41–49, 2018, doi: 10.37396/jsc.v1i1.5. [4] y. marine and s. saluky, “penerapan iot untuk kota cerdas,” itej (information technol. eng. journals), vol. 3, no. 1, pp. 36–47, 2018, doi: 10.24235/itej.v3i1.24. [5] h. yomeldi, “decision making in internet of things (iot) : a systematic literature review,” itej (information technol. eng. journals), vol. 5, no. 1, pp. 51–65, 2020, doi: 10.24235/itej.v5i1.40. [6] f. rozi, “systematic literature review pada analisis prediktif dengan iot: tren riset, metode, dan arsitektur,” j. sist. cerdas, vol. 3, no. 1, pp. 43–53, 2020, doi: 10.37396/jsc.v3i1.53. [7] a. kesuma and t. komputer, “rancangan internet of things pada kota,” vol. 2, no. 4, pp. 1–10, 2022. [8] a. g. putrada and m. abdurohman, “a systematic and critical review on activity recognition in iot-based smart lighting,” no. february, 2022. [9] a. zanella, n. bui, a. castellani, l. vangelista, and m. zorzi, “internet of things for smart cities,” ieee internet things j., vol. 1, no. 1, pp. 22–32, 2014, doi: 10.1109/jiot.2014.2306328. [10] s. n. agni, m. i. djomiy, r. fernando, and c. apriono, “evaluasi penerapan smart mobility di jakarta (evaluation of smart mobility implementation in jakarta),” j. nas. tek. elektro dan teknol. inf. |, vol. 10, no. 3, pp. 214–220, 2021. [11] syahbudin, “analisis penerapan smart city dan internet of thin gs ( iot ) di indonesia,” vol. 2016, no. 1, pp. 1–5, 2016. [12] s. hidayatulloh, “internet of things bandung smart city,” j. inform., vol. 3, no. 2, pp. 164–175, 2016. [13] e. dianawati, “future of internet of things in smart city,” itej (information technol. eng. journals), vol. 4, no. 1, pp. 39–51, 2019, doi: 10.24235/itej.v4i1.49. [14] bayu prastyo, faiz syaikhoni aziz, wahyu pribadi, and a.n. afandi, “desain banyumas smart city berbasis internet of things (iot) menggunakan fog computing architecture,” j. jeetech, vol. 1, no. 2, pp. 6–13, 2020, doi: 10.48056/jeetech.v1i2.7. [15] i. shahrour and x. xie, “kota pintar peran internet of things ( iot ) dan crowdsourcing dalam proyek smart city,” pp. 1276–1292, 2021. [16] a. rejeb, k. rejeb, s. simske, and h. treiblmaier, “internet untuk segala gambaran besar di internet halhal dan kota pintar : ulasan tentang apa yang kita ketahui dan apa yang perlu kita ketahui,” vol. 19, pp. 1–21, 2022, doi: 10.1016/j.iot.2022.100565. p-issn : 2715-2448 | e-issn : 2715-7199 vol.4 no.1 january 2023 buana information technology and computer sciences (bit and cs) 6 | vol.4 no.1, january 2023 village government readiness toward the adoption of e-government rini mayasari 1 department of informatics faculty of computer science, universitas singaperbangsa karawang rini.mayasari@staff.unsika.ac.id nono heryana 2 department of information system faculty of computer science, universitas singaperbangsa karawang nono@unsika.ac.id ‹β› agustia hananto 3 department of information system faculty of computer science, universitas buana perjuangan karawang agustia.hananto@ubpkarawang.ac.id abstrak— pelayanan publik di desa sangat penting karena desa merupakan bagian terkecil dari sistem pemerintahan administratif di indonesia dan merupakan garda terdepan dalam pelayanan publik. perkembangan teknologi mendorong pemerintah untuk mengelola pelayanan publik untuk bertransformasi memperbaiki paradigma birokrasi tidak efektif, e-government merupakan solusi untuk meningkatkan pelayanan publik. salah satu kegagalan dalam implementasi egovernment adalah ketidaksiapan pemerintah desa dalam melakukan adopsi teknologi informasi. metode penelitian yang digunakan dalam penelitian ini adalah deskriptif kualitatif. hasil yang diperoleh, model framework stope yang terdiri atas strategy, technology, organization, people dan environment merupakan model yang paling tepat untuk menganalisis kesiapan pemerintah desa. dari 5 domain yang dinilai terdapat 3 domain yang mendapatkan predikat sangat siap dan 2 domain yang mendapat predikat siap. kata kunci— e-government, stope framework, pelayanan publik abstract— public services in the village are very important because the village is the smallest part of the administrative government system in indonesia and is the front line of public services. technological developments encourage the government to manage public services to improve the ineffective bureaucratic paradigm. e-government is a solution to improve public services. one of the failures in implementing e-government is the village government's lack of readiness to adopt information technology. the research method used in this research is descriptive qualitative. the results obtained suggest that the stope framework model consisting of strategy, technology, organization, people, and environment is the most appropriate model to analyze the readiness of the village government. of the 5 domains that were assessed, 3 domains received the very ready predicate and 2 domains received the ready predicate. keywords— e-government, stope framework, public service i. introduction the village government [1] has a strategic role and position in national development because the village government is the smallest part of the administrative government system in indonesia. public services [2] in the village are a reflection of public services for higher levels of government, which so far have been carried out conventionally and have a paradigm of slow bureaucracy, convoluted procedures, and no certainty. technological developments encourage the government to manage public services to transform to improve the ineffective bureaucratic paradigm because the lack of quality public services is a major problem in indonesia. one solution is the implementation of e-government as a whole in public services. this is related to the initiation of e-government in indonesia through presidential instruction no. 6 of 2001 on telematics [3]. village government is an important level of government in a country. village government plays a crucial role in managing and developing the potential of the village [4], as well as providing public services to village residents. in the era of rapidly advancing information technology, village governments must also be ready to adopt e-government or government based on information technology. e-government is a government system that uses information technology to provide more effective, efficient, and transparent public services [5]. stope framework [6] is a framework that can be used to evaluate an organization's readiness to implement egovernment initiatives. the framework includes five components: strategy, technology, organization, people, and environment. by assessing each of these components, the organization can identify its strengths and weaknesses, and can develop a plan to overcome challenges and capitalize on opportunities. this can help the organization to successfully implement e-government initiatives and improve the delivery of services to the community[11]. technological developments encourage the government to manage public services to transform them to improve the ineffective bureaucratic paradigm. one solution is the implementation of e-government as a whole in public services. adoption of e-government [7] is absolutely necessary for now. the covid-19 pandemic forces all public services to be ready to transform into an electronicbased government system. in accordance with presidential regulation no. 95 of 2018 concerning electronic-based government systems, the implementation of e-government in mekarbuana village [8] requires various preparations to carry out the full implementation of e-government. village governments are required to be able to keep up with technological developments and continue to improve their 7 | vol.4 no.1, january 2023 ability [9] to manage population administration and public services in the village. ii. method the research method used in this research is descriptive qualitative. this is because the researcher wants to explore phenomena that cannot be quantified. as for the research framework used in this study, it is as shown in figure 1 below: fig 1. research framework iii. results and discussion based on the results of a study on the information system in mekarbuana village, there is no digital-based public information service yet. furthermore, the researchers carried out an analysis of the problems that are currently running in the process of government activities. it was found to have several problems, one of which was in carrying out various potential information systems and public information services using only conventional media. to analyze the functional requirements of the good village government system that will be made as follows: the system must be able to easily handle the process of conveying information to public acceptance. the results of literacy studies and data in the field, the required user needs related to the optimization of good village government based on information and communication technology in mekarbuana village include: 1. personal resident a. nik (population identity number) b. full name c. location of birth d. gender e. age f. religion g. phone number h. email address i. blood type j. education 2. population a. address (village, hamlet, rt, rr, street name) b. work c. resident status d. residence status 3. family a. marital status b. family card number c. father's name d. mother's name e. child relationship f. list of family members 4. etc a. passport number b. disability c. competence in addition to the need for population data collection, there are also other needs, namely data collection: 1. family, resident registration must have a family head. 2. event logging, the events recorded were residents being born, dying, moving out, and becoming tki (indonesian workers). 3. the family card separation. 4. administrative services include certificates of death; certificate of domicile; certificate of incapacity; business certificate; certificate of heirs; certificate of resident land; marriage certificate. 5. village-owned infrastructure services analysis of the implementation of good village government in mekarbuana village 1. political environment the existence of political support is the main determinant of the success of implementing good village government in mekarbuana village, with the existence of no. 95 of 2018 concerning the electronic-based government system (spbe) [10] and law (uu) no. 6 of 2014 challenging the village to make the implementation of good village government in mekarbuana village possible. 2. leadership, the commitment of the village head who shows the political will to adopt e-government to realize good village government in mekarbuana village. 3. stakeholders: there is support from various parties who have direct and indirect interests. 4. service transparency refers to the availability of information and all data that can be accessed by stakeholders and the realization of service transparency. 5. the budget, the availability of a budget to carry out the transformation of e-government. 6. technology, the presence of technology-assisted public services. results of analysis of village government readiness toward good village government in mekarbuana village using stope frameworks: table 1. stope frameworks no sub-domain percentage (%) rank (scale 4) description 1 strategy 80 4 very ready 2 technology 78 4 very ready 3 organization 75 4 very ready 4 people 61 3 ready 5 environment 60 3 ready stope 70.8 3 ready the stope framework is a tool that can be used to evaluate an organization's readiness to implement e-government initiatives. stope stands for strategy, technology, organization, people, and environment. each of these 8 | vol.4 no.1, january 2023 components is assessed on a scale from 0 to 100%, with a rank from 1 to 4. in the example provided, the organization has a very high level of readiness in the strategy and technology components, with scores of 80% and 78%, respectively. this indicates that the organization has a strong plan and direction for implementing e-government initiatives, and has a good foundation of technological infrastructure and access to technology. the organization component has a score of 75%, which is also considered very ready. this suggests that the organization has a well-defined internal structure and processes, with clear roles and responsibilities, and that there is strong support and commitment from leadership. the people component has a score of 61%, which is considered ready. this indicates that the organization's workforce is generally able to use technology effectively, but there may be some areas for improvement in terms of training and support, and in fostering a culture of engagement and participation. the environment component has a score of 60%, which is also considered ready. this suggests that there is a good level of support and participation from the community, and that there are not significant barriers in terms of legal or regulatory issues. overall, the organization has a stope score of 70.8%, which is considered ready. this indicates that the organization is well-positioned to implement e-government initiatives, but there may be some areas where additional work is needed in order to optimize the use of technology and ensure that the initiatives are successful. fig 2. village government readiness graph below are the results of a swot analysis related to the readiness of the village government toward good village government in mekarbuana village. mekarbuana village's swot analysis has an important role for mekarbuana village in realizing a good village government based on information and communication technology. swot means strengths, weaknesses, opportunities, and threats. which means strengths, weaknesses, opportunities, and threats. a. strengths the strengths of the village government's e-government initiatives can be identified by examining several key factors. one strength is that administration services are provided free of charge. this can make it easier for the community to access government services, particularly for those who may not be able to afford fees. another strength is the presence of a system that regulates administrative services. this can help to ensure that services are provided in a fair and consistent manner, and can improve the transparency and accountability of the government. another strength is the existence of a clear vision and mission that supports service quality. this can provide direction and guidance for the government as it implements e-government initiatives, and can help to ensure that services are delivered in a way that meets the needs and expectations of the community. the village government also has a strength in its management of permit application processes. by streamlining and digitizing these processes, the government can make it easier for residents and businesses to apply for permits, and can reduce the time and effort required to complete these processes. finally, the village government enjoys a high level of trust and confidence from the community. this can be a valuable asset as the government seeks to implement e-government initiatives, as it can help to encourage community engagement and participation in these initiatives. b. weaknesses there are several weaknesses that may hinder the optimization of the village government's e-government initiatives. one potential weakness is the dissemination of information to the public. if the government is not effective at communicating information about its services and initiatives, it can be difficult for the community to access and use these services. this can result in low levels of engagement and participation, and can make it difficult for the government to achieve its goals. another weakness is the current state of the online service. if the online service is not user-friendly or if it does not provide all of the necessary information and services, it can be difficult for the community to access and use it. this can limit the effectiveness of e-government initiatives and make it difficult for the government to deliver services in a timely and efficient manner. finally, the village government may still rely on manual processes and services in some cases. this can be a weakness because manual processes are typically slower, less efficient, and more error-prone than digital processes. by transitioning to digital services, the government can improve the quality and speed of its services, and can provide a better experience for the community. c. opportunities there are several opportunities for the village government to optimize its use of information and communication technology in order to improve the delivery of services to the community. one opportunity is the support of the village government. if the government is committed to investing in technology and training, and is willing to create a supportive 9 | vol.4 no.1, january 2023 and inclusive culture that encourages the use of technology, it can provide a strong foundation for the implementation of e-government initiatives. another opportunity is the high population growth rate in the village. as the population increases, the demand for government services is likely to grow as well. by using technology to improve the efficiency and effectiveness of its services, the government can meet this growing demand and ensure that the community has access to the services it needs. another opportunity is the availability of services around the clock. by providing online services and information, the government can make it easier for the community to access services at any time of day. this can improve the convenience and accessibility of government services, and can make it easier for the community to engage with the government. the advances in technology, information and communication also present an opportunity for the village government. as technology continues to evolve, the government can leverage these advances to improve the quality and speed of its services, and to provide new and innovative services that meet the needs of the community. empowering village employees through technology is another opportunity. by providing employees with the training and tools they need to use technology effectively, the government can improve the efficiency and effectiveness of its operations, and can help employees to deliver better services to the community. finally, the presidential decree no. 95 of 2018 concerning the electronic-based government system (spbe) provides an opportunity for the village government. this decree establishes a national framework for the implementation of egovernment initiatives, and can provide guidance and support to the village government as it seeks to optimize its use of technology. d. threats there are also some threats that the village government must consider in its efforts to optimize its use of information and communication technology. one potential threat is the risk of document forgery. if the government relies on digital documents and signatures, it is possible for individuals to forge these documents or signatures in order to obtain services or commit fraud. this can damage the integrity of the government and can undermine the trust of the community. another threat is the risk of malware attacks on office documents. if the government uses digital documents and files, it is possible for malware to infect these documents and cause them to be corrupted or lost. this can compromise the security and confidentiality of government information, and can disrupt the delivery of services. finally, rapid technological advancement can also be a threat to the village government. as technology continues to evolve at a rapid pace, the government must keep up with these changes in order to maintain the effectiveness and relevance of its e-government initiatives. if the government does not stay up-to-date with the latest technology, it may be left behind and unable to deliver the services that the community needs. below is a description of the results of the swot analysis related to the optimization of good village government based on information and communication technology in mekarbuana village. table 2. swot analysis strengths weakness • administration services are free of charge • system that regulates administrative services • vision & mission that supports service quality • permit application management • community trust in village government services • dissemination of information to the public. • the online service is not optimal yet. • manual service usage. opportunity threats • village government support. • high population growth rate. • service is available around the clock. • advances in technology, information and communication. • empowerment of village employees through technology. • presidential decree no. 95 of 2018 concerning the electronic-based government system (spbe). • document forgery rate • malware attack on office documents • rapid technological advancement by conducting a swot analysis, the village government can identify its strengths, weaknesses, opportunities, and threats related to the optimization of good village government based on information and communication technology. this can help the government to develop strategies to overcome challenges and capitalize on opportunities, ultimately leading to improved services and a stronger community. iv. conclusion the readiness of the village government to adopt information technology is strongly supported by various factors, either directly or indirectly. during the covid-19 pandemic, the village government was "forced" to be able to transform and innovate in providing the best service to village communities. the existing infrastructure and human resources in the village government are ready to adopt and implement technology as a whole in every public service. overall, the results of the analysis of the readiness of the village government toward good village government are in 10 | vol.4 no.1, january 2023 the rating of ready (3) on a scale of 4 to implement good village government. domain strategy has more influence on the readiness to implement e-government. references [1] s. sugiman, ‘pemerintahan desa’, binamulia hukum, vol. 7, no. 1, pp. 82–95, 2018. [2] m. h. bisri and b. t. asmoro, ‘etika pelayanan publik di indonesia’, journal of governance innovation, vol. 1, no. 1, pp. 59–76, 2019. [3] a. gioh, ‘pelayanan publik e-government di dinas komunikasi informatika kabupaten minahasa’, jurnal politico, vol. 10, no. 1, 2021. [4] n. angelia, b. m. batubara, r. zulyadi, t. w. hidayat, and r. r. hariani, ‘analysis of community institution empowerment as a village government partner in the participative development process’, budapest international research and critics institute-journal (birci-journal) vol, vol. 3, no. 2, pp. 1352–1359, 2020. [5] d. a. d. putra et al., ‘tactical steps for e-government development’, international journal of pure and applied mathematics, vol. 119, no. 15, pp. 2251–2258, 2018. [6] w. e. y. retnani, r. f. ap, and b. prasetyo, ‘analysis of user readiness level of e-government using stope framework’, in 2019 6th international conference on electrical engineering, computer science and informatics (eecsi), 2019, pp. 270–273. [7] i. k. mensah, ‘impact of government capacity and egovernment performance on the adoption of egovernment services’, international journal of public administration, 2019. [8] n. heryana, r. mayasari, a. s. y. irawan, and b. nugraha, ‘improving digital literacy skills for mekarbuana village officials’, abdimas: jurnal pengabdian masyarakat, vol. 5, no. 2, pp. 2496–2501, 2022. [9] d. mustafa, u. farida, and y. yusriadi, ‘the effectiveness of public services through e-government in makassar city’, international journal of scientific & technology research, vol. 9, no. 1, pp. 1176–1178, 2020. [10] e. prihantoro, d. mukodim, and n. r. ohorella, ‘policy analysis for the implementation of electronic-based government systems (spbe) in the city government of depok in realizing an open, participative, innovative and accountable city government’, technium soc. sci. j., vol. 19, p. 221, 2021. [11] a. alim murtopo, b. priyatna, and r. mayasari, “signature verification using the k-nearest neighbor (knn) algorithm and using the harris corner detector feature extraction method,” buana inf. technol. comput. sci. (bit cs), vol. 3, no. 2, pp. 35–40, 2022, doi: 10.36805/bit-cs.v3i2.2763. vol. 4, no.2 july 2023 | 39 iot-based farmland intrusion detection system emmanuel onwuka ibam1, olutayo k. boyinbode2, helen o. aladesiun3 1,2,3 school of computing, federal university of technology akure, nigeria eoibam@futa.edu.ng, okboyinbode@futa.edu.ng, olamipositosin@gmail.com abstract as crop vandalization with conflicts between farmers and herdsmen become recurrent in nigeria, existing farm intrusion prevention methods such as fence mounting and placement of farm guards can no longer guarantee farm security. this is because intruders either jump over the fence or attack guards on duty without visual evidence. therefore, a complementary approach using computer technologies for effective detection is required. this paper presents an iot-based farm intrusion detection model using rfid and image recognition technology. rfid sensor as well as cameras are placed at entrances of a fenced farmland for simultaneous identification. the sensor reads workers’ tags for identification, while cameras capture images of users for further identification as captured images are sent to convolutional neutral network (cnn) for recognition. a user whose image cannot be recognized is flagged as an intruder and an intrusion alert with visual evidence is sent to the farm owner. the system showed a high level of effectiveness with an accuracy of 90%, precision of 70%, and 80% recall rate and effectively controlled the rate of illegal encroachment into farmland keywords: buzzer; convolutional neural network (cnn); internet of things (iot); microcontroller; pir sensors; radio frequency identification (rfid) abstrak perusakan tanaman dengan konflik antara petani dan penggembala menjadi berulang di nigeria, metode pencegahan intrusi pertanian yang ada seperti pemasangan pagar dan penempatan penjaga pertanian tidak dapat lagi menjamin keamanan pertanian. ini karena penyusup melompati pagar atau menyerang penjaga yang bertugas tanpa bukti visual. oleh karena itu, diperlukan pendekatan pelengkap menggunakan teknologi komputer untuk deteksi yang efektif. makalah ini menyajikan model deteksi intrusi pertanian berbasis iot menggunakan teknologi rfid dan pengenalan gambar. sensor rfid serta kamera ditempatkan di pintu masuk lahan pertanian berpagar untuk identifikasi simultan. sensor membaca tag pekerja untuk identifikasi, sementara kamera mengambil gambar pengguna untuk identifikasi lebih lanjut saat gambar yang diambil dikirim ke convolutional neutral network (cnn) untuk dikenali. seorang pengguna yang gambarnya tidak dapat dikenali ditandai sebagai penyusup dan peringatan intrusi dengan bukti visual dikirim ke pemilik tambak. sistem menunjukkan tingkat efektivitas yang tinggi dengan akurasi 90%, presisi 70%, dan tingkat recall 80% dan secara efektif mengendalikan laju perambahan ilegal ke lahan pertanian. kata kunci: buzzer; jaringan syaraf konvolusional (cnn); internet of things (iot); mikrokontroler; sensor pir; identifikasi frekuensi radio (rfid) i. introduction virtually every country in the world relies on agriculture to survive, not just because it is a source of food but it is also connected to the production of most basic human needs. in nigeria, agriculture remains the leading non-oil sector of the country’s economy, providing about 70% of the nation’s population with jobs. it is a major source of livelihood for those in rural areas as they depend on the proceedings from their farm harvest to cater for their family. p-issn : 2715-2448 | e-issn : 2715-7199 vol.4 no.2 july 2023 buana information technology and computer sciences (bit and cs) vol. 2, no.2 | 40 however, the sector has suffered a major setback with persistent cases of intruders; human and animal alike causing serious damages to crops by stealing, eating, and trampling on crops. in recent times, the destruction of crops has aggravated with the incessant conflicts between herders and farm owners arising from illegal encroachment of cattle into farmlands. usually, mounting of fences round farmland, engaging farm guards, use of repellants, among others are used as means to wade off invaders. though relatively effective, these methods can be less than ideal and are sometimes prohibitively expensive to put in place [3]. moreover, with the herdersfarmers conflict, the existing security measure cannot completely guarantee safety as several farm guards have been killed and fence been jumped over without visual evidences. in order to overcome these challenges, there is need to build a system that will alert farmers of any intruder and also repel animals away from the farm. the aim of this paper is to detect and prevent intruders from entering into the farm land by implementing an iot based intrusion detection model for farmlands using rfid technology and image recognition technique. radio frequency identification technology is a technology which uses radio waves to automatically identify people or objects. currently, this technology is dramatically increasing the use of wireless technology in many areas such as healthcare, transport, military, textile, and agriculture. unlike barcode systems, rfid systems reflect better result in industrial applications like baggage tracking, access/vehicle control, animal tracking, etc. apart from traditional usage of rfid technology, another innovative research path of rfid is integrating rfid technology with mobile devices such as mobile phones and pdas [14]. this automation can provide accurate and timely information without any human intervention, access to such information where one can individually identify each one of the tagged items uniquely; help in improving your processes and also to make informed decision. radio frequency identification (rfid) technology is an automatic identification system consisting of a tag and a reader which can communicate with radio waves [11]. in this way, a lower costs more efficient process management and monitoring is provided. the remaining of this paper is organized as follows: section 2 provides the related works. section 3 describes the proposed system architecture and the methods applied to actualize the system. section 4 explains how the system was implemented, while section 5 provides the evaluation of the system. section 6 presents the conclusion. ii. related work [15] carried out a survey on animal detection methods in digital images and observed that animal in image processing have been an important field to numerous applications. they narrowed the applications to three main branches, namely detection, tracking and identification of animal. they identified the efficacy of applying computer vision in image processing for animal detection, use of transformation function such as the fourier transform, face detection approach and thresholding segmentation method. they identified lighting problem and changes of natural environment from day to night at outdoor surveillance system as two problems that need to be considered in developing an animal detection algorithm. there are copious numbers of researches undertaken to improve the old system of securing farmlands. [5], a smart farmland using raspberry pi for crop vandalization prevention and intrusion detection system was developed. they showed how passive infrared sensors (pir) was used to detect motion of human body, once the employed pir sensors detect motion the cameras capture an image and start recording the video, the owner of the farmland gets notified about the intrusion. this information along with the captured video is stored onto cloud from where the administrator /farm owner can access it once he receives the message. rfid tags was used to differentiate between the authorized person and the intruders, if the person is an authorized one then no action is taken by the system. system uses two mechanisms to ward off animals namely: the rotten egg spray and electronic vol. 4, no.2 july 2023 | 41 firecrackers. their system requires no human supervision, hence saves a lot of time and energy. the system works in real time to detect the animals in the field, in addition the farmers can access the view of their fields remotely. however, the use of fire cracker can cause damages to crops, while rotten eggs spray can result in crop diseases. [12] worked on an animal detection system using ultrasonic sensors to detect the movement of the animal and send signal. this signal is transmitted to gsm and which gives an alert to farmers and forest department immediately. however, the system can only detect the animals but cannot prevent them from destroying the crops. kaluti et al. (2018) developed an iot based wireless sensor network for easy detection and prevention of wild animal’s attack on farming lands. in their study, they showed how sensors such as pir motion sensor, sound recognizing sensor, and web cameras were used to collect data. aurdino mainly the raspberry pi, acts as master node and collects the data from the sensors which sends those data to the server for further processing. the server processes the data and sends signals to the speakers to produce sound in order to stop the animals by crossing the forest boarder, and also sends message to the mobiles of the nearer villagers, farmers, and the forest office to take the safety precautions. their system was helpful to the farmers in protecting fields and save them from financial losses and also saves them from unproductive efforts that they endure for the protection of their fields. however, in this work the stealing of rfid tags can give intruders access into the farmland. [6], worked on human-animal conflict using pir sensors and camera as first round of security where the animal movement is detected using the sensor and the sensor in turn triggers the camera to take the picture of the animal and transmit the image for processing via microcontroller i.e., through wsn. the microcontroller transmits the image from the camera to the pc in the command center where the image processing and classification of animal is done. once the animal is found to be a threat the pc will send the signal to the repellent system via microcontroller to take appropriate action. however, the system results in crop diseases. in cherukat et al., (2014), a farm field protection using sensor networks, due to the fact that fields nearer to the forests were facing problem of attack of wild animals on the crops. to this effect, they designed a smart field that is on low cost, low energy consuming, small sensor nodes. they showed how sensor nodes are deployed in groups, this group of nodes is connected to a common node and common node to the main node. the three parts of sensor deployment is primary, secondary and tertiary. the primary node controls the network, secondary nodes passes data from tertiary to primary node. by including more sensor nodes more farm areas can be covered. except the node in primary all other nodes are connected with pir (passive infrared) sensor. their system helped to protect farmlands from wild animals. it was a self-managing, cost effective and energy efficient system. however, any failure in one of the nodes results in failure of the whole system. koik et al., (2016), in their research, did a comprehensive analysis on animal detection methods in digital images. the study was carried out in order to design a system that uses digital image processing for animal detection. they showed how digital image processing system was built up by the use of power spectral in trying to test animal presence in the image, fourier transform by transforming from spatial domain to frequency domain, animal detection based on thresholding segmentation method in which if the threshold is greater than a pixel of gray that value is set to white and others are set to black. the system helped to detect the exact animal that entered the field. however, this work cannot be suitable for fast detection of animals. [8] developed an intruder recognition in a farm through wireless sensor network, this came as a result of the struggle farm owners go through for top yield in varied ways after which, their yield may be curtailed due to the interference of animals and unauthorized humans. due to this, many farmers sleep in field area to save their crops risking their lives if wild animals attack their fields. animal attack on the crops could also cause infections to the buyer when the crops are sold in the market due to the vol. 2, no.2 | 42 animal poison. hence, it is much essential to monitor the boundaries of the farm to discover movement of unauthorized entries into the farm. they therefore aimed at designing a wsn in border surveillance and intrusions detection that is cheaper for the farmer. they showed how the system is implemented to detect intrusion of animals in farms using wireless sensors and buzzers which detects the animals and produce acoustic sounds. at various locations around the farm, motion sensors are placed where certain distance is maintained between them and one of the motion sensors is made as the centralized from where we can operate all other sensors. the sensors which are present frequently sense the movement and pass it to the coordinator through rfid. an arduino board is placed near the centralized sensor to which gsm module is interfaced along with buzzers and rfid transmitter. animals are being detected by the motion sensors in the agricultural area. when an animal or human is being detected by the sensors in the agricultural area, the sensors are activated through rfid transmitter and the system produces sounds through the buzzer and will give a very minor shock to the animal. this sound irritates the animals and they cannot accommodate it at that place and due to minor shock animals will fall. however, the system was not able to provide video processing. [2], animals from wild area were continuously attacking crops for so many years and the protection of these crops field from wild animals was a serious issue. the wild animals face shortage of water and food as a result of which they move towards the agriculture area which creates great loss to the crops and annual income of farmers. when wild animals enter in a farm there is a need for an alert system to prevent crops from being damaged by wild animals. the developed a system and their objective was to prevent the loss of crops and protect the area from intrusion of wild animals which causes major damage to the agricultural area. their system has a wireless sensor network (wsn) consisting of a large number of autonomous sensors to cooperatively monitor physical or environmental conditions, such as temperature, sound, vibration, pressure, motion or pollutants. the wsn consists of various clusters connected with the sink node. each cluster has number of sensor nodes having one master node capable of collecting the data from remaining nodes, web camera and gsm connected to the raspberry pi kit. camera is used to detect the motion of wild animal and ones it gets detected it captures its image and distinguishes its features as dangerous or not, if it is dangerous then it sends instant message to the farmer. it saves farmers from unproductive efforts that they endure for the protection of their fields. however, the system is expensive and require high maintenance. [1], implemented an intelligent security system for farm protection from wild animals in order to provide food requirements of the people and produce several raw materials for industries as animal interference in agricultural lands has resulted into huge loss of crops. hence, they developed a prohibitive fencing to the farm, to avoid losses due to animals. they showed how fencing wire is used as a sensor. when animals come in contact with this open cable the circuit will be grounded and we get initial input signal that indicates presence of animals at fencing. after getting that initial input signal followed by amplifier circuit passing it for further processing, then, it will be given to the microcontroller and their system will be activated, immediately buzzer will be on, at night time, flash light will be on and message will be sent to the farmer. continuous monitoring can be done because it works on solar panel. however, damages still occurred during storms, thunder and lightning which is still a risk of dangerous shocks to farmers. therefore, the development of an internet of things farmland intrusion detection system with automatic facial recognition module for effective security and safety of both crops and farm owners aimed at tackling the shortcomings pointed in the reviewed literatures has become imperative. iii. methods based on the proposed system architecture in figure 1, the rfid reader, camera, and buzzer are all connected to a microcontroller placed at the entrance of the farmland. the rfid tags unique identification number is configured with the staff information stored in the database. at the entry point vol. 4, no.2 july 2023 | 43 into the farm, the tag is brought closer to the reader for staff identification as the staff tag is crosschecked with the staff information in the database. while confirming staff id, the microcontroller will activate the camera to capture facial image of staff. the image is then fed into face recognition module for further prove of identification. this process is carried out simultaneously to prevent intruders using authorized tag into the farmland. if any of these actions should fail, an intruder alert message is sent to the administrator’s system or the farmer’s smart phone, and the buzzer will produce a loud irritation noise. figure 1. iot-based farmland intrusion detection system architecture. the the detection of an authorized person with a valid rfid is estimated using the frequency of signal reading received by the reader. the signal power measured at the receiving node (rfid reader) is known as rssi (received signal strength indicator), the rssi received is translated from reader’s antenna into frequency of signal reading which is displayed by the reader. the rssi transmission power for the person is expressed mathematically as: 𝑝(𝑅) = 𝑝(𝑇) − 10𝑛log( 𝑅 𝑇 ) − 𝐵 (1) 𝐵 = { 𝑞 ∗ 𝑝 𝑖𝑓 𝑞 < 𝑐 𝑐 ∗ 𝑝 𝑜𝑡ℎ𝑒𝑟𝑤𝑖𝑠𝑒 (2) where n is the attenuation factor, p(t) is the signal power at the reference distance t, r is the distance between the transmitter and the receiver for the tag, q is the number of obstacles between the transmitter and the receiver, p is the attenuation factor of the farm, and c is the maximum number of obstacles between the transmitter and the receiver. the signal strength is converted into frequency of signal from the reader’s antenna and displayed by the reader. the transmission frequency increases with proximity of the rfid reader with the reference tag. the signal power in rssi is converted to distance by using the euclidean equation which is stated as: 𝐸 = √∑ (𝐴𝑛 − 𝐵𝑛)2𝑁 𝑛=1 (3) vol. 2, no.2 | 44 where e is the relative position of the reference tag a and unknown tag b, a_n represents the signal strength of the reference tag, b_n is the signal strength of the unknown tag received on the reader, n is the number of times the measurement is taken. however, the microcontroller activates the camera to take snap shot of the person. captured images were preprocessed by way of image value normalization and feature extraction. image value normalization involves bringing images into a range of into a range of intensity value that is normal. this is achieved using equation (4) 𝑜𝑢𝑡𝑝𝑢𝑡𝑐ℎ𝑎𝑛𝑛𝑒𝑙 = 255 ∗ (𝑖𝑛𝑝𝑢𝑡𝑐ℎ𝑎𝑛𝑛𝑒𝑙−𝐼𝑚𝑖𝑛) (𝐼𝑚𝑎𝑥−𝐼𝑚𝑖𝑛) (4) where 〖output〗_(channel )is the normalized image, input_channel is the image to be normalized, i_max is the maximum pixel value, and i_min is the minimum pixel value. thereafter, global features were extracted from normalized image using principal component analysis (pca). pca involves computing the eigenvalues λ and eigenvectors μ of the data correlation matrix c =x^t x, where x is an n*q data matrix of n number of samples and z features using equation (5): (𝐶 − 𝐼𝜆𝑖)𝜇𝑖 = 0 | 𝑖 = 1,2,3, … , 𝑞 (5) where q is the total number of eigenvalues. the k numbers of eigenvectors μ having the largest eigenvalues λ were picked as principal components (pcs) based on a given threshold value as defined in equation (6) ∑ 𝜆𝑖 𝑘 𝑖=1 ∑ 𝜆𝑖 𝑞 𝑖=1 ∗ 100 > 𝑡ℎ𝑟𝑒𝑠ℎ𝑜𝑙𝑑 (6) where threshold value specifies the percentage of information to be retained, k∈q represents the number of eigenvalues whose corresponding eigenvectors to be retained. pca extracted features were fed into cnn model for recognition. cnn takes the preprocessed input image vector x=(x_1, x_2,x_3,…,x_n ) with an assigned class c_l which are the images of farm staff. at the convolutional layer, images are convolved to extract features that forms feature maps f_m using equation (7) 𝑓𝑚= 𝑡𝑎𝑛ℎ(𝑤𝑥𝑖 + 𝑏) (7) where b represents the bias term, w represents weight matrix, and x_i is the input vector. thereafter, the feature maps f_m is passed to the max-pool layer where max-pooling operation is applied to each feature map to obtain the most significant features by selecting features with the maximum value as expressed in equation (8) 𝑓𝑚 = max {𝑓𝑚} (8) where f_m represents the downsized features. the obtained features f_m is fed to the fully connected layer which contain the softmax function that will classify the image data as intrusive or nonintrusive as in the equation 9. 𝑦 = 𝑠𝑜𝑓𝑡 max(𝑤𝑜𝑓𝑚 + 𝑏𝑜) (9) vol. 4, no.2 july 2023 | 45 where y represents the output, w_o represents the output weight, and b_o represent the output bias. if the recognition model give a recognition rate below 70%, it triggers an irritation alarm and subsequently send alert to the administration’s system (base station) and the owner’s smart phone iv. results and discussions this section presents the implementation of the internet of things farm land intrusion detection system, using rfid technology and image recognition technique to prevent unauthorized person into the farmland, and a safe defensive mechanism using irritation noise to scare away intruders. rfid reader was mounted at the entrance of a farm to read the tags of farm workers for identification purpose. the microcontroller, the rfid reader, tags, the buzzer, the camera and application software which includes the user interface design. the tools used in the development of the software part are as follows: • personal computer with 64 bit operating system and processor intel® of 2.20ghz. • php, mysql, html, css and javascript programming languages. • xampp apache local server. • python programing language for the image recognition algorithm. a. user interface design a user interface is the portion of a program with which a user interacts with the computer system. the user interface of the farm land intrusion detection system contains: 1. login page this is the default page that appears when the application is launched. this page enables the farm owner to access its features. figure 2 shows the system interface. the system contains three major parts, farmers login, farmers face identity prediction, farm worker enrolment. they are shown below:can be presented with tables or figures. results and discussions must also interconnect with theory that used. avoid excessive use of citations and discussion of published literature figure 2. user interface design vol. 2, no.2 | 46 figure 3. signing in of a farm owner 2. farmer’s face identification prediction the image identity checking module contains an image upload module for uploading an image which would be used to check if the face is authorized to enter into the farm. the module activates when an object or a person attempts to enter into the farmland via the entrance points. this is to prevent the possibility of an intruder using a worker’s tag. figure 5 shows the identity checking module. figure 5. identity checking interface 3. farm worker enrolment the system takes the required data and save in the database, the system wait for the farm owners to click on proceed and then put the rfid tag close to the rfid reader. figure 6 shows the unique rfid tag number and the login details to be provided by the farm owners. vol. 4, no.2 july 2023 | 47 figure 6. registration of a farm owner figure 7 shows the unique rfid tag number and the login details to be provided by the farm owners: figure 7. registration of a farm owner figure 8 shows the data of the farm owner has been successfully uploaded. vol. 2, no.2 | 48 figure 8. successful of a farm owner 4. face recognition with camera captured images will undergo different pre-process steps such as image value normalization, image enhancement, and feature extraction. figure 9. sample of pre-processed images vol. 4, no.2 july 2023 | 49 figure 10. scree plot for pca these are the output of the pre-processed images figure 11. pca reduced images figure 12 shows the face recognition operations indicating authorized and unauthorized farm owners. vol. 2, no.2 | 50 figure 12. face recognition operation figure 14 shows a layout of the system database showing farm owners that is been recognized by the system. figure 14. a layout of the system database vol. 4, no.2 july 2023 | 51 figure 15 shows typical messages sent to the farm owner’s mobile phone figure 15. sample of received message on farm owner’s mobile phone 5. evaluation confusion matrix table (shown in table 1) was used in this study to describe the performance of the recognition module on a set of test data for which the true values are known. in this study, 20 image test data were used of which 15 were authorized farm workers and 5 were regarded as intruders. the true positive (tp) represents the number of correctly detected intruders. true negative (tn) represents number of correctly recognized farmers. false positive (fp) represents the number of farmers that were incorrectly recognized as intruders, while false negative (fn) represents intruders that were incorrectly recognized as farmers. table 1. confusion matrix of system recognition module prediction class intruder farmer actual class intruder 4 1 farmer 2 14 𝐴𝑐𝑐𝑢𝑟𝑎𝑐𝑦 = 4 + 13 4 + 12 + 2 + 1 = 0.9 𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛 = 4 4+2 = 0.7 vol. 2, no.2 | 52 recall = 4 4+1 = 0.8 the system shows a high level of effectiveness with an accuracy of 90%, precision of 70%, and 80% recall rate. 6. system performance and evaluation table 2 shows the comparison of recognition accuracy score of the developed system with existing farm intrusion detection systems, while table 4 depicts the comparison of the developed system with existing system using other metrics such as efficiency, technology used, operations performed by the system and platform used. table 2. comparison of recognition rate of developed system with existing systems performance metrics vinaya et al (2018) saieshwar et al. (2018) sachin et al., (2017 developed system. accuracy 87.35% 54.32% 82.5% 90% figure 17 is the chart showing the accuracy of the previous work and developed work. figure 17. accuracy chart v. conclusion this research has established an effective farmland intrusion detection system using rfid technology and image recognition technique to prevent unauthorized person into the farmland, and a safe defensive mechanism using irritation noise to scare away intruders. rfid reader was mounted at the entrance of a farm to read the tags of farm workers for identification purpose, therefore, the system proceeded by taking face images of the workers for further identification. the capture images were 0 0,1 0,2 0,3 0,4 0,5 0,6 0,7 0,8 0,9 1 vinaya a (2018) saieshwar (2018) sachin (2017) developed research system. vol. 4, no.2 july 2023 | 53 preprocessed and salient features were extracted using pca algorithm. extracted features were fed into the recognition module driven by cnn for recognition. the system was implemented using programming language, and other packages such as flask 1.0, pytorch, mysql. the result of the system shows a recognition score of 90% accuracy, precision of 70% and 80% recall rate, while existing system shows the accuracy of 87.35%, 54.32% and 82.5% respectively. hence, with the objective of developing an iot farmland intrusion detection system using sensors and cnn, this research therefore has succeeded in controlling the rate of illegal encroachment into farmer’s farmland. references [1] abhinav v. d. (2016). design and implementation of an intelligent security system for farm protection from wild animals. international journal of science and research, 5(2), 956 -959. [2] chourey s.r, amale p.a and bhawarkar n.b. (2017). iot based wireless sensor network for prevention of crops from wild animals. special issue of international journal of electronics, communication & soft computing science and engineering, 5760. [3] felemban e. (2013). advanced border intrusion detection and surveillance using wireless sensor network technology. international journal. of communications, network and system sciences, 6, 251 -259. [4] mahesh k., naveen k. and vinaya b. (2018). iot based wireless sensor network for earlier detection and prevention of wild animals attack on farming lands. international research journal of engineering and technology, 05 (03), 1933-1935. [5] pooja g. and mohmad u.b. (2016). a smart farm land using raspberry pi crop vandalization prevention and intrusion detection system. retrieved from https://docplayer.net/90275124-asmart-farmland-using-raspberry-pi-crop-vandalization-prevention-intrusion-detectionsystem.html. [6] prajna p., soujanya b.s. and divya (2018). iot-based wild animal intrusion detection system. international journal of engineering research & technology, special issues, 1-3. [7] saieshwar r. and ramanathan r. (2018). a support vector machine with gabor features for animal intrusion detection in agriculture fields. procedia computer science, 143, 493 -501. [8] santhoshi k.j and bhavana s. (2018). intruder recognition in a farm through wireless sensor network. international journal advanced research ideas and innovations in technology, 4(3), 667 -669. [9] sachin u.s. and dharmesh j. s. (2017). a practical animal detection and collision avoidance system using computer vision technique. special section on innovations in electrical and computer engineering education, ieee access, 5, 347 -358. [10] shoukath c., ganesh r.n., abdul j.s. and manoj n. (2014), farm field protection with sensor networks. international journal of engineering research & technology, 3(11), 51 -52. [11] ustundag a. and kilinc m.s. (2010). design and development of rfid based library information system. [12] vikhram.b., revathi.b., sowmiya.s., shanmugapriya.r. and pragadeeswaran.g. (2017). animal detection system in farm areas. international journal of advanced research in computer and communication engineering, 6(3), 587-591. [13] vinaya a,, ajaykumar s. c., aditya d. b., balasubramanya k.n. and natarajan s. (2018). an efficient orb based face recognition framework for human robot interaction. procedia computer science, 133, 913 -923. [14] wickramasooriya, p.m.t.a. and thuseethan, s. (2015). understanding radio frequency identification technology: usage and integration with mobile devices. [15] boon t.k. and haidi i. (2012). a literature survey on animal detection methods in digital images.international journal of future computer and communication, 1(1), 108 -120. vol. 4, no.2 july 2023 | 76 agglomerative clustering of 2022 earthquakes in north sulawesi, indonesia berton maruli siahaan1, afrioni roma rio2 1,2department of physics, faculty of mathematic and natural science, sam ratulangi university kampus unsrat, bahu, manado, north sulawesi, indonesia 1bertonsiahaan@unsrat.ac.id, 2afrioni@unsrat.ac.id abstract this paper presents a cluster analysis of earthquake data in the surrounding region of north sulawesi, indonesia. the dataset comprises seismic data recorded throughout the year 2022, obtained from the bmkg earthquake repository. a total of 211 earthquakes were included in the analysis, with a minimum magnitude threshold of 2.5 and a maximum depth of 300 km. the agglomerative clustering technique, combined with the elbow method, was employed to determine the optimal and distinct number of clusters. as a result, four unique clusters were identified. cluster 1 exhibited high magnitudes, with an average magnitude of 4.4, and shallow depths, averaging at 20 km. cluster 2 also had high magnitudes, averaging at 4.4, but deeper depths, with an average of 199 km. cluster 3 consisted of earthquakes with low magnitudes, averaging at 3.4, and shallow depths, averaging at 21 km. lastly, cluster 4 comprised earthquakes with low magnitudes, averaging at 3.4, but deeper depths, with an average of 136 km. among the 211 earthquakes, 29 were assigned to cluster 1, 39 to cluster 2, 100 to cluster 3, which had the highest population, and 43 to cluster 4. this study provides valuable insights into the clustering patterns and characteristics of earthquakes in the region, contributing to a better understanding of seismic activity in north sulawesi, indonesia. keywords: earthquake clustering; earthquakes data, machine learning; agglomerative clustering abstrak artikel ini membahas analisis klasterisasi pada data gempa bumi di sekitar sulawesi utara, indonesia. data yang digunakan adalah data gempa bumi yang terjadi selama tahun 2022 yang diperoleh dari repositori gempa bmkg. terdapat 211 gempa bumi yang dianalisis dengan batas magnitudo terendah 2,5 dan kedalaman maksimum 300 km. dalam penelitian ini, digunakan teknik klasterisasi aglomeratif dan metode elbow untuk menentukan jumlah klaster yang optimal dan unik. hasilnya, ditemukan empat klaster yang unik. klaster 1 memiliki gempa bumi dengan magnitudo tinggi rata-rata sebesar 4,4 dan kedalaman dangkal rata-rata sebesar 20 km. klaster 2 juga memiliki gempa bumi dengan magnitudo tinggi rata-rata 4,4, namun kedalaman lebih dalam rata-rata sebesar 199 km. klaster 3 terdiri dari gempa bumi dengan magnitudo rendah rata-rata 3,4 dan kedalaman dangkal rata-rata sebesar 21 km. klaster 4 terdiri dari gempa bumi dengan magnitudo rendah rata-rata 3,4, namun kedalaman lebih dalam rata-rata sebesar 136 km. dari total 211 gempa bumi yang dianalisis, terdapat 29 gempa bumi dalam klaster 1, 39 gempa bumi dalam klaster 2, 100 gempa bumi dalam klaster 3 yang memiliki populasi terbanyak, dan 43 gempa bumi dalam klaster 4. penelitian ini memberikan pemahaman yang lebih baik mengenai pola dan karakteristik gempa bumi di wilayah sulawesi utara, indonesia. kata kunci: klasterisasi gempa bumi; data gempa; machine learning: agglomerative clustering p-issn : 2715-2448 | e-issn : 2715-7199 vol.4 no.2 july 2023 buana information technology and computer sciences (bit and cs) vol. 4, no.2 july 2023 | 77 i. introduction earthquakes are natural phenomena that frequently occur in various regions around the world, including indonesia, which is known for its high seismic activity. this region is located around the pacific ocean, known as the pacific ring of fire (rof). the pacific ring of fire is a long chain of active volcanoes and tectonic structures encircling the pacific ocean. it stretches along the western coast of south and north america, passes through the aleutian islands in alaska, extends along the eastern coast of asia through new zealand, and reaches the northern coast of antarctica. the pacific ring of fire is one of the most active geological areas on earth and often experiences powerful earthquakes and volcanic eruptions. it is home to over 450 active and inactive volcanoes. most of these volcanoes are formed through the process of subduction, where dense oceanic plates collide and slide beneath lighter continental plates. material from the ocean floor melts as it enters the earth's mantle and then rises to the surface as magma. the deepest trench in the ocean, the mariana trench, is located along the western part of the pacific ring of fire. the majority of the world's largest earthquakes also occur within this ring. these earthquakes are caused by sudden movements of rocks laterally or vertically along plate boundaries. approximately 81% of the world's largest earthquakes occur within the pacific ring of fire [1]. in indonesia, the frequency of earthquakes is exceptionally high, with an average of 6,512 tectonic earthquake events per year, equivalent to 543 events per month and 18 earthquake events per day [2]. this study focuses on the clustering of earthquakes in the vicinity of north sulawesi. the data used is sourced from the earthquake repository of the meteorology, climatology, and geophysics agency (bmkg) for a one-year period in 2022 [3]. a total of 211 earthquakes have been analyzed, meeting the minimum magnitude criteria of 2.5 and a maximum depth of 300 km. the aim of this research is to identify clustering patterns that can provide deeper insights into the seismic activity in the area. clustering is an effective approach for analyzing and grouping earthquake data based on their characteristics. several studies have applied this clustering method to categorize disaster data, disaster impacts, and tsunami potentials [4]–[7]. in this study, the agglomerative clustering method is used to partition the data into closely related groups based on attribute similarities. the validation of the unique and optimal number of clusters is performed using the elbow method. the cluster analysis results reveal the presence of four unique clusters. each cluster exhibits distinct characteristics, including magnitude scale and depth. understanding the earthquake clustering patterns in the north sulawesi region can significantly contribute to disaster mitigation efforts. by identifying the specific clusters and their characteristics, authorities and disaster management agencies can gain valuable insights into the distribution and behavior of earthquakes in the area. this knowledge can aid in the development of more targeted and effective disaster preparedness plans, early warning systems, and evacuation strategies. additionally, it can inform infrastructure development and building codes to ensure resilience against seismic events. ultimately, this research serves as a valuable resource for stakeholders involved in disaster management, enabling them to make informed decisions and implement proactive measures to reduce the potential impact of earthquakes and enhance the overall resilience of the region. ii. methods this study encompassed various stages, which involved conducting a review of relevant literature, collecting earthquake data, processing the data, applying the elbow method, performing agglomerative clustering, and conducting cluster analysis (refer to fig. 1). vol. 4, no.2 july 2023 | 78 figure 1. flowchart of research methodology. a. literature review in a high-quality research study, it is crucial to have a relevant collection of literature and references as a strong foundation. therefore, the initial stage of this research involved gathering literature related to the research topic. the collected literature encompassed various important aspects pertaining to the research subject. during the initial stage, the researcher gathered information related to earthquakes. this aimed to understand the characteristics, causes, and impacts of earthquakes that are relevant to the context of this research. additionally, the researcher studied machine learning, which is the technique or algorithm used for data processing in this study. a profound understanding of machine learning is crucial as it forms the basis for the data clustering conducted in this research. furthermore, the researcher obtained an understanding of the clustering technique using the agglomerative clustering method. clustering is a method used to group data with similar characteristics into specific clusters. in the context of this research, the application of the agglomerative clustering method enables the identification of unique clustering patterns and structures within the earthquake data in the north sulawesi region b. data gathering to obtain earthquake data in north sulawesi for the year 2022, the data source used is the official website of the meteorology, climatology, and geophysics agency (bmkg), accessible at https://repogempa.bmkg.go.id/ [3]. the geographic region parameter is set to 0°n to 3°n latitude and 123°e to 126°e longitude. although the data available on this website is limited to a single month, by leveraging knowledge of the api (application programming interface), we can efficiently access and extract data using for loops and the requests module in the python programming language [8]. the process of retrieving data through the bmkg api will generate an html file that needs further processing. to convert it into a more structured format, the data needs to be parsed into a tabular form using the beautifulsoup library [9]. once the data has been successfully parsed and organized, the next step is to save it in csv format for easier further data processing. c. data processing after successfully obtaining the earthquake data in north sulawesi for the year 2022, the next step is to process and explore the data. the data processing is performed using the python programming language [8], utilizing several packages and libraries as described below: to perform numerical calculations on arrays or matrices, the numpy library is used [10]. additionally, for data analysis and processing in the form of dataframes, the pandas library is employed https://repogempa.bmkg.go.id/ vol. 4, no.2 july 2023 | 79 [11]. for visualizing graphs and plotting data distribution on maps, the matplotlib, seaborn, and plotly libraries (including the api provided by mapbox) are utilized [12]–[15]. before proceeding with the agglomerative clustering method for clustering, the selected parameters for analysis, namely the earthquake magnitude scale (m) and depth (km), undergo a data preprocessing step using the standard scaler method. the standard scaler is a widely used technique in data preprocessing that transforms the data by subtracting the mean and dividing by the standard deviation. this process ensures that the data is centered around zero with a standard deviation of 1, making it more suitable for clustering algorithms. by applying the standard scaler, the magnitude and depth values are normalized to a comparable scale, eliminating any potential biases caused by differences in their measurement units. this normalization allows for a fair comparison and accurate clustering based on the similarity of the scaled features. d. elbow method the elbow method is a visual technique utilized to determine the optimal number of clusters for clustering algorithms. it involves plotting the explained variance by each cluster against the number of clusters and observing the point of inflection, commonly referred to as the "elbow," where adding more clusters no longer significantly improves the explained variance [16]. in simpler terms, the elbow method assists in selecting the appropriate number of clusters for the given data by identifying the point at which adding more clusters does not yield substantial enhancements in the clustering outcomes. to apply the elbow method, we initially perform clustering with various numbers of clusters and plot the explained variance against the corresponding number of clusters. the elbow point on the plot represents the stage at which the explained variance begins to level off, indicating that the inclusion of additional clusters does not contribute significantly to the improvement. once the elbow point is determined, we can choose the number of clusters that strikes a balance between the explained variance and the simplicity of the model. e. agglomerative clustering once the optimal number of clusters has been determined using the elbow method, the subsequent step involves applying agglomerative clustering to the preprocessed data. agglomerative clustering is a hierarchical clustering technique that begins with each data point assigned as a separate cluster and gradually merges the closest clusters until all data points are grouped into a single cluster [17]. in this research, we utilized the agglomerative clustering algorithm provided by the python scikitlearn library [18]. the algorithm requires specifying the number of clusters, which is set to the optimal number obtained from the elbow method. for this study, ward's linkage criterion was employed, aiming to minimize the sum of squared differences within all clusters. f. cluster analysis afterwards, the resulting clusters are analyzed to identify patterns or trends within the data. this analysis involves examining the characteristics of each cluster, such as the average values of relevant variables, as well as utilizing visualizations like box plots to depict the clusters. by conducting cluster analysis, we can gain insights into the data structure and identify meaningful clusters among the population of 211 earthquakes in north sulawesi. these clusters can be further investigated to obtain a deeper understanding of the characteristics specific to each cluster. iii. results and discussion in this section, we will delve into the findings derived from the analysis of earthquake clusters in the sulawesi utara region of indonesia for the year 2022. the application of the elbow method resulted in the identification of four distinct clusters, as illustrated in fig. 2. these clusters represent groups of earthquakes that share similar characteristics in terms of their magnitudes and depths. vol. 4, no.2 july 2023 | 80 following the determination of the optimal number of clusters, the agglomerative clustering technique was employed to further examine the data. this hierarchical clustering method starts by considering each earthquake event as an individual cluster and then iteratively merges the two closest clusters until all data points are grouped into a single cluster. by applying this method to the identified four clusters, we were able to discern notable dissimilarities among them, particularly in terms of the magnitudes and depths of the earthquakes they encompass. figure 2. elbow method with agglomerative clustering out of the total of 211 earthquakes analyzed, we observed that the first cluster consisted of 29 earthquakes, the second cluster contained 39 earthquakes, the third cluster emerged as the largest group with 100 earthquakes, and the fourth cluster comprised 43 earthquakes. these cluster-specific earthquake populations can be visualized in fig. 3. figure 3. distribution of earthquake events based on cluster. vol. 4, no.2 july 2023 | 81 the distribution of earthquake data and the formed clusters can be found on the map displayed in fig. 4. in the cluster analysis, cluster 3 emerged as the most dominant cluster, primarily distributed in the offshore area. this cluster exhibits characteristics of low magnitude scale, with an average of 3.4 m, and shallow depths, with an average of 21 km. therefore, this area tends to be relatively safe from the impacts of earthquakes. figure 4. clustering results: grouping of north sulawesi earthquake events in 2022 no cluster 1 was found in the main island area. this cluster needs to be closely monitored as it exhibits a relatively high magnitude scale, with an average of 4.4 m, and shallow depths, with an average of 20 km. shallow depths can increase the potential for damage compared to deeper depths. however, in the northern island areas of mainland sulawesi, there are several earthquake sources that fall within cluster 1. therefore, this area requires careful mitigation planning to reduce the risk for the local population. vol. 4, no.2 july 2023 | 82 figure 5. distribution of magnitude based on cluster. cluster 2 also displays a high magnitude scale, with an average of approximately 4.4 magnitude, and deep depths with an average of 199 km. this type of earthquake often occurs in the main island area, which may be related to the activities of active volcanoes in the north sulawesi region. figure 6. distribution of depth based on cluster. cluster 4 exhibits a low magnitude scale and deep depths, making it relatively safe. however, it is still important to remain vigilant regarding the potential earthquakes within this cluster. the distribution and descriptive data of the clustering results can be observed in fig. 5 and fig. 6, along with table 1. vol. 4, no.2 july 2023 | 83 table 1. data description based on cluster cluster number of earthquakes average magnitude (m) average depth (km) 1 29 4.4 20 2 39 4.4 199 3 100 3.4 21 4 43 3.4 136 having an understanding of the characteristics of the formed clusters, this information can be utilized in disaster mitigation efforts and more effective planning to protect the community and reduce the risks associated with earthquake impacts in north sulawesi. by incorporating this knowledge, appropriate measures can be taken to enhance preparedness, response, and resilience in the face of seismic events. it is crucial to prioritize the safety and well-being of the population and ensure that strategies are in place to mitigate the potential consequences of earthquakes. you can access the results of the clustering analysis through the following link: https://sl.unsrat.ac.id/eq-ac. iv. conclusions in conclusion, the application of agglomerative clustering to analyze earthquake data in the sulawesi utara region resulted in the identification of four distinct clusters with varying characteristics in terms of magnitude and depth. these clusters offer valuable insights into the seismic activity patterns in the area. the findings of this study have significant implications for disaster mitigation and risk reduction efforts in sulawesi utara. by understanding the clustering patterns of earthquakes, stakeholders can develop more effective strategies to protect the local population and minimize the impact of seismic events. future research in this field could focus on expanding the dataset by incorporating data from additional sources and over a longer time period. this would provide a more comprehensive understanding of the seismic activity and clustering patterns in sulawesi utara. additionally, investigating the correlation between these earthquake clusters and geological features, such as fault lines or volcanic activity, could provide valuable insights for predicting and mitigating future seismic events. v. references [1] m. masum and m. a. akbar, “the pacific ring of fire is working as a home country of geothermal resources in the world,” in iop conference series: earth and environmental science, iop publishing, 2019, p. 012020. [2] a. sabtaji, “statistik kejadian gempa bumi tektonik tiap provinsi di wilayah indonesia selama 11 tahun pengamatan (2009-2019),” bul. meteorol. klimatol. dan geofis., vol. 1, no. 7, pp. 31–46, 2020. [3] b. m. k. dan geofisika, “eq repository,” 2023. https://repogempa.bmkg.go.id/ [4] p. novianti, d. setyorini, and u. rafflesia, “k-means cluster analysis in earthquake epicenter clustering,” int. j. adv. intell. inform., vol. 3, no. 2, pp. 81–89, 2017. [5] m. murdiaty, a. angela, and c. sylvia, “pengelompokkan data bencana alam berdasarkan wilayah, waktu, jumlah korban dan kerusakan fasilitas dengan algoritma k-means,” j. media inform. budidarma, vol. 4, no. 3, pp. 744–752, 2020. [6] m. t. furqon and l. muflikhah, “clustering the potential risk of tsunami using density-based spatial clustering of application with noise (dbscan),” j. environ. eng. sustain. technol., vol. 3, no. 1, pp. 1–8, 2016. https://sl.unsrat.ac.id/eq-ac vol. 4, no.2 july 2023 | 84 [7] a. wahyu and r. rushendra, “klasterisasi dampak bencana gempa bumi menggunakan algoritma k-means di pulau jawa,” jepin j. edukasi dan penelit. inform., vol. 8, no. 1, pp. 174–179, 2022. [8] m. f. sanner and others, “python: a programming language for software integration and development,” j mol graph model, vol. 17, no. 1, pp. 57–61, 1999. [9] l. richardson, “beautiful soup documentation.” april, 2007. [10] c. r. harris et al., “array programming with numpy,” nature, vol. 585, no. 7825, pp. 357–362, 2020. [11] w. mckinney and others, “pandas: a foundational python library for data analysis and statistics,” python high perform. sci. comput., vol. 14, no. 9, pp. 1–9, 2011. [12] p. barrett, j. hunter, j. t. miller, j.-c. hsu, and p. greenfield, “matplotlib–a portable python plotting package,” in astronomical data analysis software and systems xiv, 2005, p. 91. [13] m. l. waskom, “seaborn: statistical data visualization,” j. open source softw., vol. 6, no. 60, p. 3021, 2021. [14] p. t. inc, “collaborative data science,” 2015. https://plot.ly [15] mapbox, “map, geocoding and navigations apis at: https://docs.mapbox.com/api/maps/.” accessed, 2023. [16] p. bholowalia and a. kumar, “ebk-means: a clustering technique based on elbow method and k-means in wsn,” int. j. comput. appl., vol. 105, no. 9, 2014. [17] d. müllner, “modern hierarchical, agglomerative clustering algorithms,” arxiv prepr. arxiv11092378, 2011. [18] f. pedregosa et al., “scikit-learn: machine learning in python,” j. mach. learn. res., vol. 12, pp. 2825–2830, 2011. p-issn : 2715-2448 | e-issn : 2715-7199 vol.4 no.1 january 2023 buana information technology and computer sciences (bit and cs) 19 | vol.4 no.1, january 2023 critical path method in contractor service company management information systems using incremental model hamim tohari 1 study program text accounting department of computerized accounting, politeknik negeri madiun email: htohari@pnm.ac.id ahmad kudhori 2 study program text accounting department of computerized accounting, politeknik negeri madiun email: akudhori@pnm.ac.id ‹β› hedi pandowo 3 study program text accounting department of computerized accounting, politeknik negeri madiun email: hedipandowo@pnm.ac.id abstract— good control is needed in managing a project, starting from controlling the human resources to systematic scheduling. cv. xyz surabaya is a contractor service company, that does not yet have an information system that can be used as a control tool in project management. critical path method (cpm) is one method that can be used to schedule projects. this study aims to design a project management information system by implementing cpm. the system design is done using the incremental model. the results of this study are in the form of a prototype system that fits the needs of cv. xyz surabaya, includes system flow, contextual data model (cdm), and user interface (ui). key words—cpm, information system, incremental abstrak— pengendalian yang baik diperlukan dalam mengelola suatu proyek, mulai dari pengendalian sumber daya manusia hingga penjadwalan yang sistematis. cv. xyz surabaya merupakan perusahaan jasa kontraktor yang belum memiliki sistem informasi yang dapat digunakan sebagai alat kontrol dalam manajemen proyek. critical path method (cpm) merupakan salah satu metode yang dapat digunakan untuk menjadwalkan proyek. penelitian ini bertujuan untuk merancang sistem informasi manajemen proyek dengan mengimplementasikan cpm. perancangan sistem dilakukan dengan menggunakan incremental model. hasil dari penelitian ini berupa sistem prototype yang sesuai dengan kebutuhan cv. xyz surabaya, meliputi alur sistem, contextual data model (cdm), dan user interface (ui). kata kunci—cpm, sistem informasi, incremental i. introduction good control is needed in the management of a project. the control starts from controlling human resources to structured scheduling, and other factors that affect the progress of the project. in addition to influencing the progress of the implementation of a project, these factors can also be the cause of the minimum delay in project completion, so that the planned time does not exceed the predetermined time. if a project has a problem, it will have an impact on the implementation of the project, if the implementation of a project fails then the goals that have been set also fail, and will cause a waste of time and costs. cv. xyz surabaya is a company engaged in construction services with a specialization in floor construction, which includes water proofing, floor hardener, sealant polyurethane, concrete additive, bonding agent, and concrete injection, counting, epoxy floor & wall, termite control, cutter & ripere services. concrete. the absence of a good management information system (according to the needs of cv. xyz surabaya) has an impact on business processes in project management that are not well organized. cpm is one method that can be used to plan and supervise projects, which is the method most widely used by many systems that use the approach to the principle of network formation. the use of the cpm method can save time in completing various stages of a project [1]. observing the existence of problems in project management in cv. xyz surabaya, a solution that might be used is to create a web-based contractor management information system that will implement cpm in the project planning and scheduling process. in more detail, this contractor management information system is expected to assist in planning, minimize the occurrence of discrepancies in the plan and project realization, and facilitate the process of paying workers and filling out report documents. cpm is an activity-oriented activity that schedules project activities through network drawing [2]. levin and kirkpatrick in [1] explain that cpm(critical path method) is a method for planning and supervising projects which is the most widely used system among all other systems that use the principle of network formation. the cpm method is widely used by industry or construction projects. this method can be used if the duration of the work can be known and is not too fluctuating. siswanto in [1] explains that cpm is a project management model that prioritizes cost as the object being analyzed, cpm is a network analysis that seeks to optimize the total project cost by reducing the total project completion time. mailto:htohari@pnm.ac.id mailto:akudhori@pnm.ac.id 20 | vol.4 no.1, january 2023 we can know the critical path by calculating two start and end times for each activity, 1) early start (es), which is the previous time an activity can start, assuming all predecessors have finished. 2) the earliest finish (ef), which is the previous time activity could be completed. 3) last start (ls), which is the last time an activity can be started so as not to delay the completion time of the entire project. 4) the last finish (lf), which is the last time an activity can be completed so that it does not delay the completion time of the entire project (see figure 1). es a ef ls d lf figure 1. critical path elements description: a = activity name d = the duration of an activity es = earliest start ls = latest start ef = earliest finish lf = latest finish slack time is the free time that each activity has to be able to be postponed without causing delays in the overall project. slack time can be formulated as follows: 𝒔𝒍𝒂𝒄𝒌 = 𝑳𝑺 − 𝑬𝑺 or 𝒔𝒍𝒂𝒄𝒌 = 𝑳𝑭 − 𝑬𝑭 description: slack = free time ls = latest start es = earliest start lf = latest finish ef = earliest finish according to larman in [3], it is stated that the iterative model is a methodology that relies on the development of software applications one step at a time in the form of expanding the model. this methodology is based on the initial specification of the basic model of the application being built. according to k. schwalbe in [4], the incremental model is "the incremental build life cycle model provides for the progressive development of operational software, with each release providing added capabilities". the incremental process model uses repetitive linear sequences to build software. as time goes by, each linear sequence will result in developments in software work which can then be used by users [5]. each stage in the increment method, which is contained in the methodology, has input and output. the output of the increment process will be used as input for the next increment process [6]. the incremental model was chosen because it has several advantages, namely: an easy process, there is testing and debugging, the possibility of project failure is small, and it can produce software according to needs in a relatively shorter time [7]. the incremental model is a method consisting of several increments with simple management, where the product is designed, implemented, and tested in stages (each module will be added in stages) until the product is declared complete or as needed. an information system is a system within an organization, which brings together the needs of daily transaction processing that can support the operational functions of a managerial organization with the strategic activities of an organization that can provide reports needed by certain outside parties [8]. an information system is an organized combination of people, hardware, software, communication networks, and databases that collect, transform and disseminate information in an organizational form [9]. according to [10] state that "management information system is a computer-based system that makes information available to users who have similar needs". referring to some of these references, it can be stated that the management information system is a structured set of elements that can present the information needed by management to support decision-making. this set of elements includes people, hardware, software, databases, and procedures. project management is a process of planning, organizing, and controlling company resources with short-term goals to achieve objectives and specific goals. project management is designed to manage and control company resources according to related activities, time efficiency, cost efficiency, and good performance. this requires good processing and can be achieved. some of the things that need to be managed in the project management area include cost, quality, occupational health and safety, environmental resources, risk, and information systems. project management is the application of knowledge, expertise, and skills, the best technical methods, and with limited resources, to achieve predetermined goals and objectives to obtain optimal results in terms of cost performance, quality and time, and work safety [11]. the project management process can be concluded as shown in figure 2. figure 2. project management process it takes a contractor who can carry out the work of the owner so that the work (project) can be carried out as planned. a contractor or can be referred to as a contractor, is a person or a business entity that is bound by a project and carries out the project by the contract agreement that has been made/agreed upon. a contractor is a person or entity who accepts work and carries out the implementation of the work according to the costs that have been determined based on the plans and regulations, as well as the conditions that have been set [12]. a contractor can be declared a person or an institution who is responsible for working on a project according to the specifications and budget provided by the owner or project provider. 21 | vol.4 no.1, january 2023 ii. method this type of research is research & development (r&d). the information system design process is carried out using an incremental model.jenis penelitian ini adalah research & development (r&d). research design the design and construction activities of the contractor service company's information system are carried out using the incremental model. thus, the research design was made to follow or adapt to the existing stages in the incremental model, by the established boundaries(see figure 3). figure 3. research design iii. results and discussion 1. system requirements analysis new system design required by cv. xyz surabaya will involve five users with different access rights. the five users consist of admin, owner, project manager, customer, and foreman. the processes contained in the project management system in cv. xyz surabaya consists of (1) worker data input process, (2) service data input process, (3) equipment data input process, (4) material data input process, (5) user data input process, (6) order input process, (7) project data input process, (8) transaction process, (9) cpm calculation process, (10) project scheduling process, (11) attendance process, (12) progress input process, and (13) salary management process. input data user sistem admin start input nama, username, password, hak akses p h a s e end daftar user mst_user figure 4. flow map user data input process input pesanan pelanggan sistem admin start input data pelanggan, data pesanan pesanan prencatatan pesanan daftar pesanan buat surat penawaran kirim surat penawaran daftar pesanan end p h a s e figure 5. flow map order input process input data proyek sistem pemilikadmin p h a s e pesanan_detail prencatatan data proyek pilih acc pesanan, input biaya, durasi, tanggal mulai, tanggla selesai, unggah dokumen pendukung daftar data proyek start daftar data proyek end figure 6. flow map project data input process 22 | vol.4 no.1, january 2023 hitung cpm sistem manajer proyek start input nama kegiatan, alat yg digunakan, pekerja, durasi kegiatan, kegiatan pendukung pesanan_kerja_ alat pesanan_kerja_ bahan pesanan_kerja_ kegiatan pesanan_kerja_ke giatan_pekerja perhitungan es (early start) ls (last start) ef(early finish) lf(last finish) perhitungan waktu bebas/slack (s) s=ls-es atau s=lf-ef end tabel cpm p h a s e pembuatan/input jalur kritis figure 7. cpm calculation process 2. contextual data model(cdm) the cdm design which will then be used to design the application of the system is shown in figure 8. figure 8. contextual data model(cdm) design 3. user interface(ui) design based on the results of the analysis of system requirements (functional requirements), ui designs are obtained for each process consisting of ui for (1) material master pages, (2) tools master pages, (3) user data pages, (4) data pages workers, (5) customer order master page, (6) transaction data page, (7) wage data page, (8) customer page, (9) service listing page, (10) payment page, (11) payment details page, (12) project data page, (13) worker attendance data page, (14) login page as shown in figure 5, (15) service list page, (16) service list page, (17) worker attendance master page, (18) attendance details page, (19) project progress list page, (20) customer order report page, (21) transaction report page, (22) salary report page, (23) schedule monitoring page. all of these uis are integrated into the main ui, which is the home page as shown in figure 10. figure 9. login page figure 10. home page in line with the research by [4] entitled “implementation of the incremental model in the information system for leasing of goods and services pt. sriwijaya indah persada palembang". the results of his research stated that information systems can be built using an incremental model consisting of requirements, specifications, architecture design, code, and test stages, while this study resulted in system design, cdm, and ui. while the research by [13] entitled "development of patient data information system for the rehabilitation section of bnn malang city using iterative incremental method", is also relevant to this study, in which the results of this study state that the results of the development of the information system are carried out according to the needs, this is evident from the results of the scenario testing carried out, while the results of this study also show that there is a conformity between the results of the system design and the needs reviewed based on the results of the system requirements analysis. iv. conclusion referring to the results and discussion that have been previously discussed, and after being associated with the formulation of the problem that has been determined in this study, the following conclusions can be drawn; the process of designing a contractor service company management information system using the incremental model can produce a prototype system that includes a system flow map, cdm, and ui. the application of the cpm method in the design of the management information system of the contractor service company at cv. xyz surabaya is by system requirements. upah_pekerja pekerja_absesnsi absen_detail report_kegiatan alat_kerja pekaerja_jasa kerja_bahan kegiatan_pekerja gaji_pesanan pesan_proses kerja_kegiatan pesan pesan_jasa bayar_pesan absensi id_absensi tanggal id_kegiatan id_do keterangan batal cruser crtime mduser mdtime variable characters (15) date integer variable characters (15) variable characters (250) bitmap (1) variable characters (5) date & time variable characters (5) date & time identifier_1 absensi_detail id id_pekerja presensi catatan id_absensi integer variable characters (5) boolean variable characters (250) variable characters (15) identifier_1 log_data id operation tabel user waktu old_value new_value integer variable characters (15) variable characters (150) variable characters (10) timestamp blob blob mst_alat id_alat nama_alat keterangan gambar status_aktif cruser crtime mduser mdtime variable characters (15) variable characters (60) variable characters (250) variable characters (100) boolean variable characters (5) date & time variable characters (5) date & time identifier_1 mst_bahan id_bahan nama_bahan keterangan gambar status_aktif cruser crtime mduser mdtime variable characters (15) variable characters (60) variable characters (250) variable characters (100) boolean variable characters (5) date & time variable characters (5) date & time identifier_1 mst_jasa id_jasa nama_jasa keterangan gambar status_aktif cruser crtime mduser mdtime variable characters (15) variable characters (60) variable characters (250) variable characters (100) boolean variable characters (5) date & time variable characters (5) date & time identifier_1 mst_pekerja id_pekerja id_jasa ava nama_lengkap upah_harian npwp tempat_lahir tanggal_lahir jenis_kelamin ktp alamat_sekarang alamat_ktp telp phone email status_aktif cruser crtime mduser mdtime variable characters (5) variable characters (15) variable characters (250) variable characters (200) double variable characters (60) variable characters (35) date variable characters (20) variable characters (250) variable characters (250) variable characters (50) variable characters (18) variable characters (60) boolean variable characters (5) date & time variable characters (5) date & time identifier_1 identifier_2 mst_user id_user nama_lengkap username password status_aktif hak_akses cruser crtime mduser mdtime variable characters (5) variable characters (200) variable characters (100) variable characters (50) boolean enum variable characters (5) date & time variable characters (5) date & time identifier_1 penggajian id_penggajian id_order tanggal keterangan batal cruser crtime mduser mdtime variable characters (15) variable characters (15) date variable characters (250) bitmap (1) variable characters (5) date & time variable characters (5) date & time identifier_1 penggajian_detail id id_penggajian id_pekerja bayar upah integer variable characters (15) variable characters (5) integer double identifier_1 pesanan id_order id_pemb nama_order alamat phone telp catatan crtime variable characters (15) variable characters (20) variable characters (60) variable characters (250) variable characters (18) variable characters (50) variable characters (250) date & time identifier_1 pesanan_detail id id_order id_jasa luas lokasi jumlah integer variable characters (15) variable characters (15) double variable characters (250) double identifier_1 pesanan_bayar id_payment id_order pembayaran_ke tanggal nominal attach keterangan keterangan_admin konfirmasi cruser crtime mduser mdtime variable characters (15) variable characters (15) integer date double variable characters (100) variable characters (250) variable characters (250) enum variable characters (5) date & time variable characters (5) date & time identifier_1 pesanan_kerja id_do id_order tanggal keterangan batal kerja_selesai cruser crtime mduser mdtime variable characters (15) variable characters (15) date variable characters (250) bitmap (1) bitmap (1) variable characters (5) date & time variable characters (5) date & time identifier_1 pesanan_kerja_alat id id_do id_alat jumlah keterangan terpenuhi catatan integer variable characters (15) variable characters (15) double variable characters (250) boolean variable characters (250) identifier_1 pesanan_kerja_bahan id id_do id_bahan keterangan terpenuhi catatan jumlah integer variable characters (15) variable characters (15) variable characters (250) boolean variable characters (250) double identifier_1 pesanan_kerja_kegiatan id id_do id_jasa nama_kegiatan tgl_mulai tgl_target keterangan integer variable characters (15) variable characters (15) variable characters (100) date date variable characters (250) identifier_1 pesanan_kerja_kegiatan_pekerja id id_kegiatan id_pekerja integer integer variable characters (5) identifier_1 pesanan_proses id id_order penawaran proses gagal kontrak tgl_mulai tgl_selesai deal crtime integer variable characters (15) boolean boolean boolean boolean date date boolean date & time identifier_1 report_kegiatan id_report id_kegiatan tanggal presentase keterangan dokumentasi batal cruser crtime mduser mdtime variable characters (15) integer date integer variable characters (250) variable characters (100) bitmap (1) variable characters (5) date & time variable characters (5) date & time identifier_1 settings_cetakan id kode template kertas cruser crtime mduser mdtime integer variable characters (100) mediumtext variable characters (30) variable characters (5) date & time variable characters (5) date & time identifier_1 23 | vol.4 no.1, january 2023 acknowledgment thanks to all those who have provided assistance and support to us, especially to the politeknik negeri madiun through lp3m which has provided funding for this research. references [1] nugraha, arif rakhmat. evaluasi pelaksanaan proyek dengan metode cpm dan pert (studi kasus pembangunan terminal binuang baru kec. binuang). tugas akhir, program studi teknik industri, fakultas teknologi industri universitas islam, yogyakarta. 2016. [2] nishi sharma. applications of critical path method in project management. international journal of management and economics. vol.1.no.26. 2018. [3] budi, d.s., dkk. analisis pemilihan penerapan proyek metodologi pengembangan rekayasa perangkat lunak. teknika. vol. 5, no.1. 2016. [4] arsia, r., azdy a.r. implementasi incremental model pada sistem informasi penyewaan barang dan jasa pt. sriwijaya indah persada palembang. teknik informatika stmik palcomtech palembang. vol.6, no.2. pp. 1-9, 2016. [5] pressman, roger s. rekayasa perangkat lunak. yogyakarta: andi. 2010. [6] syarif. m., nugraha. w. metode incrmental dalam membangun aplikasi identifikasi gaya belajar untuk meningkatkan hasil belajar siswa. jusikom:jurnal sistem komputer musirawas. vol.4. no.1. pp. 43-50. 2019. [7] rather. m.a., bhatnagar. v.a. comparative study of software development life cycle models. international journal of application or innovation in engineering & management. vol. 4. no.10. p. 7. 2015. [8] sutabri, tata. analisis sistem informasi. yogyakarta: andi. 2012. [9] o'brien, james. a. introduction to information systems. new york: mcgraw-hill. 2005. [10] mcleod raymond, jr. sistem informasi manajemen. ed.10. jakarta: salemba empat. 2008. [11] husein, abrar. manajemen proyek , perencanaan, penjadwalan & pengendalian proyek. yogyakarta : andi offset. 2008. [12] ervianto, i.w. manajemen proyek konstruksi edisi revisi. yogyakarta: andi. 2005. [13] prastya s.e., saputra, m.c., pramono d. pengembangan sistem informasi data pasien seksi rehabilitasi bnn kota malang menggunakan metode iterative incremental. jurnal pengembangan teknologi informasi dan ilmu komputer. vol. 2. no.12. pp. 6587-6596. 2018. vol. 5, no.1 | 45 detection of malacca woven fabric motifs using the yolov4 method adi semri neno1, aviv yuniar rahman2, fitri marisa3 1,2,3department of informatic engineering, universitas widyagama malang, indonesia 2school graduate, doctor of philosofhy in information & communication technology, asia universiti, selangor, malaysia 1nenoady@gmail.com, 2aviv@widyagama.ac.id, 3fitri@widyagama.ac.id abstract malacca is one of the districts that has a weaving culture and also produces woven cloth in east nusa tenggara. the large number of types of woven cloth from each malacca tribe means that outsiders and even native malacca people are not yet familiar with typical malacca motifs, therefore a system is needed that can help make it easier for people to recognize the types of woven fabric motifs. malacca woven fabric in this study was used to detect the types of woven fabric motifs in malacca district using the yolov4 method. the results of detecting malacca woven fabric motifs correspond to each type of woven fabric. apart from that, the malacca woven fabric motif detection system with yolov4 technology is an effective and efficient solution in recognizing malacca woven fabric motifs. malacca woven fabric is classified into four classes with an impressive map score of 100%. keywords: object detection, identifying, malacca woven fabric motifs, woven fabric, yolov4. i. introduction woven fabrics are one of indonesia's valuable cultural heritages, exuding the rich traditions and folk arts of each region[1]. one area known for its beautiful and unique woven fabric is malacca regency, east nusa tenggara province[2]. malacca woven cloth is also an important symbol in local culture, reflecting the history, beliefs and values of its people[3]. one of the most famous cultural assets of malacca regency is the art of traditional woven cloth. malacca woven cloth is famous for its beautiful and colorful designs[4]. the woven fabric motifs often reflect the surrounding culture and nature. in addition, the practice of weaving is a skill that is passed down from generation to generation[5]. it helps develop craftsmen's skills and keeps the tradition of arts and crafts alive, as it has beautiful and colorful designs, and woven fabric motifs that reflect the culture and natural surroundings[6]. malacca woven cloth not only plays a role as traditional clothing, but also has a deeper role[7]. in everyday life. motifs and patterns resulting from traditional weaving techniques become a means of telling ancient stories, local mythology, and passing knowledge between generations. there are many types of woven cloth motifs from each malacca tribe, so outsiders and even native malacca people are not yet familiar with the typical malacca woven cloth motifs[8]. therefore, it is necessary to detect woven fabric motifs which can help make it easier for the public to recognize the type of woven fabric motif using the yolov4 method[9]. in previous research, hue, saturation, value (hsv) and gray level cooccurrence matrix (glcm) feature extraction was carried out to identify woven fabric motifs in south central timor regency. research was carried out to identify types of woven fabrics in tts district using the hsv color feature extraction method, and glcm texture characteristics, and to measure the similarity of woven fabrics using the euclidean distance metho. the p-issn: 2715-2448 | e-issn: 2715-7199 vol.5 no.1 januari 2024 buana information technology and computer sciences (bit and cs) vol.1, no.2 | 46 results obtained in this research obtained a glcm texture accuracy level for color features of 55%, hsv color features of 62.5% and combination of color and texture features of 91.67%[10]. the aim of this research is to detect malacca woven cloth motifs using yolov4, so that it can help foreigners and native malacca people to recognize the types of motifs on malacca woven cloth[11]. this research will make a positive contribution to preserving culture, education and economic development in malacca regency, as well as introducing the beauty of malacca woven cloth art to the wider world through a modern technological approach[12]. ii. methods in this case, it is a method for detecting malacca woven fabric motifs using yolov4. the methodology used in this research is shown in figure 1. it begins with the first process of literature study which will be carried out by researchers to look for references for implementing malacca woven fabric motif detection using yolov4. this process is carried out by running a script that has been designed by the researcher. then testing was carried out using a dataset prepared in the form of woven fabric. the test results will then be evaluated using the mean average precision parameter. the purpose of this evaluation is to compare the results of identifying yolov4 objects in malacca woven fabric motifs. fig 1. motifs detection flow malacca woven fabric using yolov4 1. literature review in this case the researcher looked for references from several sources related to yolo. in this process, researchers also use references within the limits of only using the yolo method. on the other hand, researchers also used source journals to look for references in detecting malacca woven fabric motifs. 2. data collection the source that has been obtained is the data used in the reference for yolo object detection of various types. the results of the collection will later be implemented into malacca woven fabric motifs yolov4. 3. pre-processing in this case, pre-processing is a process of classifying woven fabric motifs according to type and class. the way to classify woven fabric motifs is to create a bounding box. where these limits, the box will later become a parameter in producing output, namely according to the class of woven fabric motif. 4. mhetod implementation the application method used in this process is yolov4. the researchers designed code to implement yolov4 for woven fabric motifs. this will later be executed according to each code that has been designed by the researcher. 5. yolo method training the training process in the yolov4 method is the process of running the code designed by the researcher. the training process involves five categories of malacca woven fabric motifs motig_garuda, vol.1, no.2 | 47 motif_marobo futus, motif_human and deer, motif_futus men dataset will be used to test the results of this training. 6. yolo method testing the implementation method used in this process is using yolov4. researchers designed code to implement yolov4 for detection. this will later be executed according to each code that has been designed by the researcher. 7. evaluation the final step is the evaluation stage, which involves assessing the results obtained from the yolov4 small test. in testing, the parameters used for evaluation include mean average precision (map), which is calculated based on the equation. map = 1 𝑛 ∑ 𝑖 𝑛 =1 api [13] in the equation, “n” represents the actual value, while “i” corresponds to the curve value on the precision x and y axes. the resulting plot will involve point interpolation to separate the resulting curve from the x and y axes. iii. result and discussion table i is the results of the tests carried out. this could explain that starting from 1000 iterations, the garuda motif map value is 82.10% of the total between the training data and testing data. then for the marobo futus motif, mpa results were obtained at 100% in the detection of malacca woven fabric motif objects. the male futus motif produces an map level of 97.65% for object detection in malacca woven fabric motifs. furthermore, testing on human and deer motives, the final test resulted in an map score of 65.58%, a fairly large difference between the data used for training. to find out to what extent yolov4's detection accuracy is accurate, the testing process continues until the 6000th iteration. there are differences or discrepancies between the data used for training and testing at the end of the evaluation. table 1. resulut from woven fabric motifs type iteration 1000 2000 3000 4000 5000 6000 eagle motif 82.10% 100% 100% 100% 100% 100% marobo futus motifs 100% 100% 100% 100% 100% 100% humen and deer motifs 65.58% 100% 100% 100% 100% 100% men’s futus motifs 97.65% 100% 100% 100% 100% 100% furthermore, in the 2000 iteration, the garuda motif already had a map value of 100% of the total difference or distinction between the data used for training and testing purposes. then, the marobo futus motif also has an map result of 100% in detecting malacca woven fabric motifs. the male futus motif produces a map level of 100% detection of the malacca woven fabric motif object. furthermore, testing on human and deer motifs had an map level of 100% of the total detection. the next test used 3000 iterations, the results of the garuda motif had a map value of 100% of the total sum of the differences between the data used for training and testing purposes. then the marobo futus motif has a result of 100% in the detection of malacca woven fabric motif objects. the male futu motif produces a map detection rate of 100%. woven fabric motif objects. furthermore, testing on human and deer motifs alone had the same map level. in detecting other motifs, the map level obtained was 100% of the total detected. in the 2000th to the 6000th iteration, the map level obtained was more stable and did not experience a decrease. various iterations have maximum yields of map levels up to 100% vol.1, no.2 | 48 fig 2. map highest detection of malacca woven fabric motifs fig 3. detection results of malacca woven fabric motifs table 2. comparison of research methods with the proposed method researcher object detection method map fs lesiangi ay mauko and, bs djahi 1). tts woven fabric 2).3 fabric images tts tribal weaving datasets 1). hue, saturation, value hsv), dan 55%,-62,5% 2). gray level cooccurrence matrix (glcm) 91,67%. our proposal woven fabric 1500 datasets yolov4 100.0% in figure 2 it can be explained that the values produced in the tenu malaka fabric motif detection test with various iterations had maximum results with map levels reaching 100.0%. for the time used in detection is only a few minutes. however, if detection with more iterations, it will take around 8 hours to produce mean accuracy precision. additionally, at the maximum batch used in testing, 14,000 sampling tests and training data were used. the error rate in the entire test was only 0.522 in the woven fabric motif detection test. the tests that have been carried out have the maximum and highest map values from 1000 iterations to 6000 iterations shown in figure 3. it can be explained that the map at 1000 iterations is the test results has a maximum value of 97.65%. for the highest map, namely 2000 iterations up to 6000 iterations, the maximum result is 100%. in figure 3, it can be explained that the detection of malacca woven fabric motifs using yolov4 has been successful and can detect the types of malacca woven fabric motifs according to each class. the tests carried out to detect yolov4 objects are also very short and efficient in terms of time and accuracy. therefore, the yolov4 method is very effective in detecting various motifs of malacca vol.1, no.2 | 49 woven fabric. the results obtained by the yolov4 method can help the outside community to recognize malacca woven fabric motifs. in table ii there is a comparison between the methods that have been used to detect malacca woven fabric motifs with the proposed method. previous research identified woven fabric motifs using the hue, saturation, value (hsv) and gray level cooccurrence matrix (glcm) methods. this research used 3 images of tts tribal woven fabric with 2 methods studied. the results of this research were the highest, namely 91.67% in identifying tts woven fabric using the gray level coocrrence matrix method. this result is relatively high, but in this case the research only used data on 3 images of tts tribal woven fabric. the data is said to be very small because of the large number of woven fabrics. from the proposed goal, the researchers used malacca woven fabric motifs using 1500 data with 4 types of classes for the process of detecting malacca woven fabric motifs. the testing process uses 1500 motig image data of malacca woven fabric. the results obtained from this test were higher than the previous method, namely 100% detected using the yolov4 method on malacca woven fabric motifs. iv. conclusions the results of detecting malacca woven fabric motifs using the yolov4 method prove that detecting malacca woven fabric motifs according to the maximum class of malacca woven fabric motifs, namely men's futus motifs, and the 1000th iteration produces an map level of 97.65%. and produced high map in the 4 classes of malacca woven fabric motifs from 2000 iterations to 6000 iterations with the highest map of 100%. this result is the very best result. references [1] l. d. e. koten, r. safitri, and m. p. wulandari, “hermeneutics of ikat weaving (utan) lian lipa from sikka regency, east nusa tenggara (ntt),” int. j. sci. soc., vol. 3, no. 3, pp. 107– 118, 2021, doi: 10.54783/ijsoc.v3i3.358. [2] y. nataliani, “frieze group in generating traditional cloth motifs of the east nusa tenggara province,” jtam (jurnal teor. dan apl. mat., vol. 6, no. 3, p. 651, 2022, doi: 10.31764/jtam.v6i3.8568. [3] s. soetrisno, d. r. sulistyaningrum, and i. bifawa’idati, “texture-based woven image classification using fuzzy c-means algorithm,” int. j. comput. sci. appl. math., vol. 8, no. 1, p. 1, 2022, doi: 10.12962/j24775401.v8i1.9588. [4] y. suryati and b. n. nggarang, “analysis of working postures on the low back pain incidence in traditional songket weaving craftsmen in ketang manggarai village, ntt,” j. epidemiol. public heal., vol. 5, no. 4, pp. 469–476, 2020, doi: 10.26911/jepublichealth.2020.05.04.09. [5] j. l. . bessie, a. s. w. langga, and d. n. s. sunbanu, “analysis of marketing strategies in dealing with business competition (study on ruba muri ikat weaving msme in kupang city),” webology, vol. 19, no. 1, pp. 780–794, 2022, doi: 10.14704/web/v19i1/web19055. [6] h. hambali, m. mahayadi, and ..., “classification of lombok songket cloth image using convolution neural network method (cnn),” pilar nusa mandiri …, no. 85, pp. 149–156, 2021, doi: 10.33480/pilar.v17i2.2705. [7] t. r. s. hidayat, nurindah, and d. a. sunarto, “developing of indonesian colored cotton varieties to support sustainable traditional woven fabric industry,” iop conf. ser. earth environ. sci., vol. 418, no. 1, 2020, doi: 10.1088/1755-1315/418/1/012073. [8] j. s. nalenan, f. siki, m. r. talan, and r. j. wabang, “cultural ideology in woven fabric motif of insana communities at the indonesian– timor leste border,” int. j. lang. cult., vol. 3, no. 2, pp. 57–63, 2021. [9] x. long et al., “pp-yolo: an effective and efficient implementation of object detector,” 2020, [online]. available: http://arxiv.org/abs/2007.12099 [10] f. s. lesiangi, a. y. mauko, and b. s. djahi, “feature extraction hue, saturation, value (hsv) vol.1, no.2 | 50 and gray level cooccurrence matrix (glcm) for identification of woven fabric motifs in south central timor regency,” j. phys. conf. ser., vol. 2017, no. 1, 2021, doi: 10.1088/17426596/2017/1/012010. [11] e. halim et al., “the application of digital module design of east sumba woven fabric on interior accessories,” pp. 286–294, 2022, doi: 10.5220/0010752600003113. [12] j. terven and d. cordova-esparza, “a comprehensive review of yolo: from yolov1 and beyond,” pp. 1–34, 2023, [online]. available: http://arxiv.org/abs/2304.00501 [13] s. lestari, “ieequality detection of export purple sweet potatoes using yolov4-tiny”. vol. 5, no.1 | 39 detection of diseases and pests on the leaves of sweet potato plants sing yolov4 melita srinosdian nisti1, aviv yuniar rahman2, fitri marisa3 1 2 3departement of informatic engineering, universitas widyagama malang, indonesia jl. borobudur no.35, mojolangu, kec.lowokwaru, kota malang, jawa timur 2 school graduate, doctor of philosofy in information & communication technology, asia university, selangor, malaysiah 1 melitasrinosdiannisti@gmail.com, 2 aviv@widyagama.ac.id, 3 fitri@widyagama.ac.id abstract sweet potato (ipomea batats) is a root plant that can live in all weather, in mountainous areas and on the coast. this plant is one of the important food crops in indonesia, and makes indonesia the second largest sweet potato producer after china. however, according to data from the central statistics agency (bps), sweet potato production in indonesia in 2018 decreased by 5.63% when compared to production in 2017 which reached 1,914,244 tons [2]. based on these data, it is important to conduct research on pest and disease detection in plants. therefore, the author conducted a study related to this problem entitled detection of diseases and pests on the leaves of sweet potato plants using yolov4 with the aim of helping educate farmers in recognizing diseases on the leaves of sweet potato plants and how to overcome them. in this study the dataset was sweet potato leaves with a total of 1500 data divided into three classes, namely aspidomorpha, yellow spot and normal leaves with 4000 iterations. the best training results on 1500 data with 75% accuracy. the yolov4 algorithm produces high accuracy in detecting diseases in the leaves of sweet potato plants. keywords: yolov4, detection, sweet potato, plants i. introduction sweet potatoes are vines that live in all weather conditions, mountainous areas and coastal areas. sweet potatoes are a very nutritious food with quite a lot of carbohydrates and calories. therefore, sweet potatoes are also used as a staple food in about 4,444 regions. sweet potatoes are also a good source of vitamins and minerals [4]. sweet potato (ipomoea batatas) is one of the important food crops in indonesia and makes indonesia the second largest sweet potato producer in the world after china. indonesia's sweet potato production in 2018 was 1,806,389 tons of tubers. sweet potato production decreased by 5.63% compared to 2017 production of 1,914,244 tons. in addition to the decline in sweet potato production, the harvest area has also decreased. the sweet potato harvest area in indonesia in 2018 was 90,707 ha, a decrease of 14.61% compared to the 2017 harvest area of 106,266 ha (ministry of agriculture of the republic of indonesia, 2019). previous research on leaf disease in plants, namely potatoes. in addition to agriculture, much has been done in the field of technology in overcoming the problem of potato leaf disease, one of which is the use of information technology in detecting potato plant diseases through image processing or commonly called digital image handling [10]. you only look once (yolo) is an algorithm developed for real-time object detection. the detection system used must use a recycling classifier or locator for its detection. the model is applied to images at various locations and scales. observation is considered the area where the image scores highest [5]. yolo uses a very different approach, applying a single neural network to the entire image. this network divides the image into several regions, then predicts the bounding box and its probability, for each box in the boundary region, there is a possibility to classify it as an object or not [9]. this research in its application is detecting diseases in sweet potato leaves using the primary dataset obtained by itself from cv. raj organic, with categories in one dau getting one image divided into three classes, namely: 1. aspidomorpha, 2. yellow spot, 3.and normal. the purpose of this study p-issn: 2715-2448 | e-issn: 2715-7199 vol.5 no.1 januari 2024 buana information technology and computer sciences (bit and cs) mailto:melitasrinosdiannisti@gmail.com mailto:aviv@widyagama.ac.id mailto:fitri@widyagama.ac.id vol.5, no.1 | 40 is to educate farmers in recognizing and understanding diseases and pests on plant leaves and testing the accuracy of the yolov4 method in detecting. leaf diseases in this study include: 1. yellow spot yellow spots on plant leaves can be a sign of a variety of problems, including disease, nutrient deficiencies, or environmental stress. to identify the exact cause of yellow spoton plant leaves, it is necessary to examine in more detail and consider factors such as plant type, growth environment, and other symptoms that may be associated. here are some possible causes of yellow spoton plant leaves: such as nutritional deficiencies especially nitrogen [8]. nitrogen-deficient plants will experience yellow leaves, especially on the underside of the plant but many also occur on the surface of the leaves themselves. 2. aspidomorpha aspidomorpha is a genus of a group of insects known as leaf beetles. these insects are often found on various types of plants, including sweet potato plants. they are insects that feed on leaves and can become pests for plants if their population is too large. this if allowed to eat can cause bare leaves and affect the quality of sweet potatoes produced [4]. ii. methods figure 1. flowchart stages of research (source: personal preparation) 1. data collection the image used in this study is the image of sweet potato leaves. with the number of 1500 dataset images with jpg format, which are divided into several classes, namely normal leaves 500, yellow spot 500 and aspidomorpha 500 data. the data collection technique is to use primary data obtained directly from the subjects of the first perpetrators of a study. taken using a samsung galaxy a10s cellphone. 2. data preprocessing the things done in research related to object detection include literature studies where the author looks for references and problems that will be researched. then conducted interviews in the field in this study located at raj organic, sukun malang interviewed directly the founder of cv raj organik regarding the problem to be researched. this technique is more efficient in obtaining information about the leaf disease of sweet potato plants. furthermore, the collection of image data taken using a cellphone camera with the total amount of data is 1500 which is further divided into 3 classes according to the type of plant disease. next is the process of melting and changing the size of the image. overview is the initial term in which a dataset is labeled for the purpose of storing picture information. this process is done by giving a bounding box with a class name to each image object. then changes the image size is carried out to enhance the performance of the yolo example in object recognition [3]. vol.5, no.1 | 41 3. yolo configuration network configuration is required as a model network to load the data to be trained. the configuration of parameters used for the formation of the yolov4 model is adjusted to the number of object classes to be detected and the ability of the graphics card to handle computational processes owned by the researcher [6]. the parameters involved are batch with value 64 and subdivision with value 16. this parameter means that in one process the step is able to computate 64 data with the division of 16 data. 4. testing testing is done by selecting the model that provides the best accuracy at the end of the training process. after obtaining a model that has good and stable performance will be loaded again in the testing process. next, the model will be given input images to start its testing process. tests carried out on the yolov4 algorithm will provide detection results in the image in the form of marker boxes and object labels based on their class. when the training process runs it will produce the best weight where this best weight will be used to detect objects captured by the camera. after that, the display will display information and information from the object where the information obtained is in the form of confidence values [3]. 5. analysis at this stage, the process of calculating the performance of the testing stages on each algorithm is also carried out. researchers analyzed the mean average precision or map. the analysis method was chosen to compare the actual number of objects and the number of correct objects detected. iii. results and discussions the test was conducted with 4 different epoch stages. the number of datasets is 1500 data with epochs of 1000,2000,3000 and 4000. in a detection system, the ground station (laptop) must be able to detect leaf imagery using yolov4. the stage carried out is to collect datasets in the form of images with a total of 15000 images of various types of sweet potato leaf diseases, then labeling. 1. establishment of detection model modeling is a step after the training process is carried out to form a model or weight recognition system using yolov4 recognition objects [12]. after each image is labeled which is training data, the next step is training. in this study, we used yolov4 or a higher-weighted version. the training process creates a weight expansion file that is used to detect sweet potato leaf objects. figure 2. training dataset result source: (personal preparation) vol.5, no.1 | 42 based on the results of the training conducted to create the weight file took 86 hours with 4000 iterations of epochs. in figure 5, the highest accuracy is 75%. 2. test results the training result data is in the form of weights, cfgs, and file names used by the jalar ui leaf detection system. the system that has been created can be run in real time using a web camera. its detection uses programs built with the python language yolov4. tests conducted as shown in figures 5 and 6 show that the system is able to detect sweet potato leaves according to the class in real time by providing a bounding box when clearly identified [7]. 3. confusion matrix evaluation table 1. confusion matrix results details, precision, recall, f-score. (source: personal preparation) precision and recall are two calculations commonly used to measure system performance. in this study, precision is calculated to determine the degree of precision between the requested information and the system's response, recall is the extent to which the system has succeeded in finding information, precision is the predicted value and the actual value [1]. based on the table above, the 2000 iteration has a high precision of 0.75% and the lowest precision of 0.12%. it can also be seen in the recall graph that the high value is at 1000 iterations with a percentage of 1.00% and the lowest recall is 0.30%. the f-score contrasts with the weighted average of precision and memory. according to the aforementioned findings, the system does not work well in categorizing positive responses of each type, as shown by poor precision and recall results, which also indicate a low true positive (tp) score. higher precision and recall values because they have higher true positive (tp) values [11]. however, because the number of false positives (fp) is also high, causing the f-score value to be less than optimal. here is the f-score chart. 4. system testing system testing is performed to demonstrate system performance. system tests are also run to check whether the system recognizes objects correctly. when testing, the system displays a bounding box that contains confidence values. the value obtained can change due to lack of lighting on the object or taking a picture that is less clear so that the object is blurred. below are the results of testing the object system by its class. itera si t p fp f n presi si reca ll fscor e 1000 1 2 5 0 0.70 % 1.00 % 0.82 % 2000 1 2 4 1 0.75 % 0.92 % 0.82 % 3000 5 1 37 2 9 7 0.12 % 0.34 % 0.18 % 4000 3 6 16 2 8 4 0.18 % 0.30 % 0.23 % vol.5, no.1 | 43 table 2. confidence values (source: personal preparation) object confidence value information 0.52 normal leaves 0.30 aspidomorphause the "insert citation" button to add citations to this document. 0.37 yellow spots 0.87 normal leaves based on the test results above, fish objects can be recognized well and the system can display information including confidence values. the confidence value changes due to changes in the contrast of light on the object, but according to the test results, the confidence value is higher when the distance between the leaf object and the adjacent camera and the leaf object is photographed according to its class hasl bounding also affects the value of confidance (widjaja &; leonesta, 2022. iv. conclusions according to research findings, system testing requires a reliable internet connection and less signal interference. based on the results of testing and image processing with yolov4, the system can detect fires in real time with 4000 training epochs, thus achieving the highest accuracy of 75%. . the yolov4 test also obtained an map accuracy of 73.3%, with the best precision of 0.75%, recall of 1.00%, f-score of 0.82%. this study showed good performance in its tests. vol.5, no.1 | 44 references [1] ahmad, t., ma, y., yahya, m., ahmad, b., nazir, s., haq, a. u., &; ali, r. (2020). object detection through modified yolo neural network. scientific programming, 2020, 1–10. https://doi.org/10.1155/2020/8403262 [2] gultom. (2021). chapter i حض خ ِ ي. galang cape, 2504, 1–9. [3] hardiansyah, b., &; primasetya, a. (2023). face mask detection system using yolov4 deep learning algorithm. stains (national seminar on technology & science), 2(1), 313–318. [4] hondo, a., damaiyanti, k. c., hafizh, m. f., amanatillah, n. e., &; damayanti, t. a. (2018). sweet potato yellow spot disease in bogor, west java. indonesian journal of phytopathology, 14(2), 69. https://doi.org/10.14692/jfi.14.2.69 [5] k, k. c. (2023). plant disease detection and classification using deep learning. international journal for research in applied science and engineering technology, 11(5), 6305–6308. https://doi.org/10.22214/ijraset.2023.52103 [6] lestari, s. (n.d.). ieequality detection of export purple sweet potatoes using yolov4-tiny. [7] li, c., li, l., jiang, h., weng, k., geng, y., li, l., ke, z., li, q., cheng, m., nie, w., li, y., zhang, b., liang, y., zhou, l., xu, x., chu, x., wei, x., &; wei, x. (2022). yolov6: a singlestage object detection framework for industrial applications. http://arxiv.org/abs/2209.02976 [8] pradana, a. w., samiyarsih, s., &; muljowati, j. s. (2017). correlation of anatomical characters of sweet potato leaves the cultivar is resistant and does not withstand the intensity of leaf scurvy. scripta biologica, 4(1), 21. https://doi.org/10.20884/1.sb.2017.4.1.381 [9] prasanta, m. r., &; pranata, m. y. (2021). fire detection system design using yolov4 framework. https://semnastera.polteksmi.ac.id/index.php/semnastera/article/view/243 [10] rozaqi, a. j., sunyoto, a., &; arief, m. rudyanto. (2021). disease detection in potato leaves using image processing with convolutional neural network method. creative information technology journal, 8(1), 22. https://doi.org/10.24076/citec.2021v8i1.263 [11] sandy eka putra, casi setianingsih, &; ratna astuti nugrahaeni. (2022). detection of parking violations on the shoulder of toll roads with the intelligent transportation system using ssd algorithm. e-proceedings of engineering, 9(3), 1064–1069. [12] sauqi, m. (2022). vehicle detection using you only look once (yolo) v3 algorithm. islamic university of indonesia, 5–8. [13] widjaja, p. a., &; leonesta, j. r. (2022). determining mango plant types using yolov4. formosa journal of science and technology, 1(8), 1143–1150. https://doi.org/10.55927/fjst.v1i8.2155 vol.5, no.1 | 45 vol. 4, no.2 july 2023 | 85 j.valarmathi1, v.t.kruthika2 liver disease prediction model based on oversampling dataset with rfe feature selection using ann and adaboost algorithms ahmed sami jaddoa1, samah j. saba2, elaf a.abd al-kareem3 1 business informatics college, university of information technology and communications, iraq 2 department of computer science, science of college, university of diyala, iraq 3 department of sharia, college of islamic sciences, university of diyala, iraq ahmed.sami@uoitc.edu.iq 1, samah.j.saba@gmail.com 2, elaaf.ali1989@gmail.com 3 abstract liver disease counts are one of the most prevalent diseases all over the world and they are becoming very common these days and can be dangerous. liver diseases are increasing all over the world due to different factors such as excess alcohol consumption, drinking contaminated water, eating contaminated food, and exposure to polluted air. the liver is involved in many functions related to the human body and if not functioned properly can affect the other parts too. predication of the disease at an earlier stage can help reduce the risk of severity. this paper implemented oversampling dataset, feature selecting attributes, and performance analysis for the improvement of the accuracy of classification of liver patients in 3 phases. in the first phase, the z-score normalization algorithm has been implemented to the original liver patient data-sets that has been collected from the uci repository and then works on oversampling the balanced dataset. in the second phase, feature selection of attributes is more important by using rfe feature selection. in the third phase, classification algorithms are applied to the data-set. finally, evaluation has been performed based upon the values of accuracy. thus, outputs shown from proposed classification implementations indicate that ann algorithm performs better than adaboost algorithm with the help of feature selection with a 92.77% accuracy. keywords: machine learning, classification, feature selection, rfe, ann, adaboost, and liver. abstrak hitungan penyakit hati adalah salah satu penyakit yang paling umum di seluruh dunia dan menjadi sangat umum akhir-akhir ini dan bisa berbahaya. penyakit hati meningkat di seluruh dunia karena berbagai faktor seperti konsumsi alkohol berlebihan, minum air yang terkontaminasi, makan makanan yang terkontaminasi, dan paparan udara yang tercemar. hati terlibat dalam banyak fungsi yang berkaitan dengan tubuh manusia dan jika tidak berfungsi dengan baik dapat mempengaruhi bagian lain juga. predikasi penyakit pada tahap awal dapat membantu mengurangi risiko keparahan. makalah ini mengimplementasikan dataset oversampling, atribut pemilihan fitur, dan analisis kinerja untuk peningkatan akurasi klasifikasi pasien hati dalam 3 fase. pada tahap pertama, algoritme normalisasi zscore telah diimplementasikan ke kumpulan data pasien hati asli yang telah dikumpulkan dari repositori uci dan kemudian bekerja pada oversampling kumpulan data yang seimbang. pada tahap kedua, pemilihan fitur atribut lebih penting dengan menggunakan pemilihan fitur rfe. pada fase ketiga, p-issn : 2715-2448 | e-issn : 2715-7199 vol.4 no.2 july 2023 buana information technology and computer sciences (bit and cs) vol. 4, no.2 july 2023 | 86 algoritma klasifikasi diterapkan pada kumpulan data. akhirnya, evaluasi telah dilakukan berdasarkan nilai-nilai akurasi. dengan demikian, keluaran yang ditunjukkan dari implementasi klasifikasi yang diusulkan menunjukkan bahwa algoritma jst memiliki kinerja yang lebih baik daripada algoritma adaboost dengan bantuan pemilihan fitur dengan akurasi 92,77%. kata kunci: pembelajaran mesin, klasifikasi, pemilihan fitur, rfe, ann, adaboost, dan liver i. introduction liver disease can be defined as liver inflammation that results from the actions of bacteria, or toxic materials so that liver doesn’t properly operate anymore. according to the reports that have been conducted by world health organization (who) 2005 there has been an estimate that 7.6 million patients had died from cancer and 84 million individuals would die over the next decade. this data had shown that the liver cancer represents 6th most widespread cancer type worldwide and it is the 3rd-largest death cause along with the development. it’s unavoidable that technology development and easier access to internet have made it easier to identify liver disease and become big supporters of dealing with special need illnesses [1]. machine learning (ml) represents an artificial intelligence (ai) part that allows the system to get knowledge without any explicit knowledge. the supervised algorithms take advantage of the human inputs and outputs for prediction accuracy and training process, which is why, they are utilized for a variety of the applications of classification. thus, ml application had extended to the health-care also. a very significant problem in the health-care is the rising numbers of the liver disease patients. liver is one of the most vital organs with some functionalities such as detoxification of chemicals, bile production, and productions of vital protein types for the blood clotting [2].feature selection has also been referred to as the instance selection, attribute selection, variable selection, data selection, feature construction, or feature extraction. it is utilized for the data reduction by redundant and removing irrelevant data for increasing data mining accuracy. feature selection chooses many relevant features from original features [3].classification has been defined as one of the crucial tasks in dm and ml, due to the fact that it is aimed at categorizing every instance in the dataset to distinctive groups on the basis of information that has been identified by its features. in addition to that, a major dm task is the data classification. it has been attempted to create classifier identifying diabetes at minimal cost and with optimal performance [4][5]. ii. literature review over the recent years, various researches have been performed to classify liver patients. s. jain et al. [6], proposed a paper based on the indian liver patient dataset that has a variety of symptoms for around 600 patients. this work is aimed at the evaluation of several intelligent technique outputs, such as k-nn, xgboost, support vector machines (svm), and decision tree with the ratio of the training set to testing set being 80% and 20% respectively. and results have shown that k-nn gives an accuracy of 64%, the svm model gives a 66% accuracy, the decision tree model gives an 81% accuracy, and xgboost gives an accuracy of 91%. g. jamila et al. [7], the proposed model for the prediction of liver cirrhosis sickness employed naive bayesian, classification and regression tree (cart), and svms with 10-fold cross-validation. accuracy, recall, precision, and f1 score were used for the evaluations of the model's performance. vol. 4, no.2 july 2023 | 87 among all the strategies used in this study, svm technique produces the optimal results, with an accuracy of 73%, precision of 73%, recall of 100%, and f1 score of 84%. g. s. harshpreet kaur [8], this study has been based upon the prediction of the liver diseases with the use of ml algorithms. the prediction of the liver diseases involves many different levels of steps, such as: preprocessing, classification and feature extraction. in this paper, a hybrid classification approach has been suggested for the prediction of liver diseases, and data-sets have been collected from kaggle data-base of indian liver patient records. the suggested model was able to achieve a 77.58% accuracy. m. ghosh et al. [9], aimed at evaluating a number of the ml outputs, such as random forest, logistic regression, xgboost, svms, adaboost, decision tree and k-nn for prediction and diagnosis of the chronic liver disease. the algorithms of the classification have been assessed on the basis of different criteria of measurement, like the accuracy, f1 score, precision, recall, area under the curve (auc), and specificity. amongst algorithms, random forest exhibited superior performance in the prediction of liver diseases with 83.7% accuracy. n. nahar et al. [10], analyzed a new and efficient method of ensemble learning for classification of liver diseases, where 5 ensemble algorithms, namely adaboost, beggrep, logitboost, begg-j48, and random forest have been implemented and compared based on accuracy, fpr, rmse tpr, and roc curve. logitboost outperformed the rest of the ensemble methods, where its accuracy has been 71.53%. this paper has codified an effective process for diagnosing liver disease using deep learning giving it a web-based approach. the model attained an accuracy of 67.6 percent and this model predicts whether the user is having a liver disease or not. iii. materials and method fig.1. overall process of liver disease model a. dataset and attributes presently, there is a wide range of the data-sets related to liver diseases. in the present paper, ilpd has been utilized, it includes 583 rows and 2 classes. where 1st class is associated with the patient records (prs) of the liver disease and includes 416 records, the 2nd one is for the non-liver (pr) and consists of 167 records determined with the use of the summation of every sector field. fig1 illustrates the distribution of the data in data-set. in general, the data-set includes 11 columns for 142 females and 441 male patients. details have been listed in table1. vol. 4, no.2 july 2023 | 88 table1. attributes of the dataset no attributes type range 1 age: patient age interval [4-90] 2 gender: patient gender nominal [femalemale] 3 tb: total bilirubin interval [0.40-75] 4 db: direct bilirubin interval [0.10-19.70] 5 alkphos: alkaline phosphotase interval [63-2,110] 6 sgpt alamine: amino-transferase interval [10-2,000] 7 sgot aspartate: amino-transferase interval [10-4,929] 8 tp: total protiens interval [2.70-9.60] 9 alb: albumin interval [0.90-5.50] 10 a/g ratio: ratio of albumin and globulin interval [0.30-2.80] 11 selector field * binary [1-2] fig. 1. the number of patients in the dataset b. dataset pre-processing pre-processing can be defined as a highly vital stage in ml classification as the cleaner the data, then the better are the result of classification tends to be [11]. the methods of preprocessing that have been applied in the model can be explained as: a. reducing noisy data: there are 2 data noise types in ml, which include: class noise and attribute noise. none-the-less, for the maximum accuracy in suggested model, the attribute noise is decreased for enhanced accuracy with the use of panda library. b. data transformation: which indicates the process of the reorganization or re-structuring of the raw data. it’s utilized for the purpose of transforming the raw data to proper format allowing the data mining to obtain the strategic information faster and in a more effective way. c. standard scalar: which transforms the data in a way that its distribution has an average value of 0 as well as a standard deviation that equals to 1. the aggregate functions conduct the operations on column values them return one value. c. oversampling oversampling refers to the random duplication of the minority class values. as we have already seen, the ipld dataset has 167 non-liver samples and 416 liver samples. therefore, it may suffer from imbalanced class distribution issue that the class of the majority may bias prediction. to overcome this problem, random oversampling is used to increase the majority of class samples. d. feature scaling feature scaling normalizes feature values in a pre-defined range. it’s a very vital step for building a machine learning model. it reduces the training time and sometimes helps to achieve faster 1 6 7 4 1 6 n o n l i v e r l i v e r vol. 4, no.2 july 2023 | 89 convergence for many machines learning. scaling using mean and standard deviation may suffer if a dataset contains too many outliers. we have used the z-score outlier detection technique to detect the outliers and handle those outliers using robust scaling. e. feature selection feature selection can be defined as the process of the selection of significant characteristics strongly associated with output from data-set for faster model training, decreased dimensionality, reduced complexity, improved accuracy and straightforward interpretation. significant bio-markers/variables have been obtained from records of clinical information and lab tests of the patients with the use of the ml and statistical data mining algorithms. the abovementioned preprocessing tools include packages allowing feature selection [12]. recursive feature elimination (rfe) rfe can be defined as feature selection approach of a wrapper type. internally, it utilizes filter-based approaches; none-the-less, it differs from filter method. it has 2 significant options of configuration, which include: i. it determines the number of the features that are to be chosen, ii. it sets ml algorithm in the feature selection. in initial case, it searches a sub-set of the features through the consideration of all of the features that are present in training data-set and eliminates features until the needed number of the features is left. in 2nd case, it utilizes an ml algorithm and ranks characteristics based on their significance. it discards least significant features then repeats model fitting steps. the entire process is repeated to the point where the stated number of the features is left [13]. f. dataset splitting the data-set is split to data for process analysis training and testing. in that, 80% of the data has been utilized for the training and 20% of it has been used for the testing. g. classification techniques classification is a model used to predict the future behavior of the data by classifying the records into predefined classes. in the classification, precise disease detection with the use of the testing and training dataset [14]. it proposed 2 ml models for building prediction. initially, the training data has been trained across 2 ml models, such as neural network and adaboost are predicted based on a trained model of learning, one by one, and after that, test the data. some of the parameters that include the precision, accuracy, and recall are finally compared with some algorithms that have been explained above. a. artificial neural network classifier an ann [15] is a simulation of the working of the biological neural networks. each one of the nodes has been modeled after a neuron, which is why, it is referred to as artificial neurons as well. an nn is made up of several layers, every one of which has a number of the nodes. the typical nn has been represented by fig2. vol. 4, no.2 july 2023 | 90 fig2. diagram of a typical ann basically, there are 3 components in the typical ann: • input layer – one layer whose number of the nodes is dependent upon the input dimensions. the input layer applies a transform to nn’s input and passes that along as input to hidden layers. • output layer which is the last layer of an nn, the dimensions of which have been characterized by the output. this layer conducts a functionality on hidden layer’s output prior to the production of the results. • hidden layer those layers represent the algorithm’s crux. they conduct all of the calculations on input for the purpose of producing output. the work of those layers is not known. which is why, only weights and parameters that have been provided to those layers may be tweaked for the purpose of producing the needed results. a network becomes deeper with the increase of the number of the hidden layers. each one of the nodes in a network is referred to as a perceptron, which has been depicted in fig3. a perceptron is made up of 2 parts, which are: a sum of inputs and activation function on summation. a certain node takes weighted summation of its inputs then passes it to linear or nonlinear activation function. fig3. a diagram of the perceptron the equation for certain perceptron has been depicted by 1. weighted summation of inputs (x.w) is passed through activation function (f) besides bias value (b). it may be denoted as product of vector dot, where n represents the number of the inputs for each node. the activation function produces output prediction that has been provided as set of the inputs. bias term has been added to computation for the purpose of helping in the enhancement of the learning of the perceptron. z = f (b + x.w) = f (b + ∑ 𝑥𝑖 𝑤𝑖 𝑛 𝑖=1 ) (1) each perceptron utilizes step function as the activation function. in a set of the perceptron’s, which is ann (referred to as the multi-layer perceptron as well), each one of the layers may have a separate activation function. vol. 4, no.2 july 2023 | 91 b. adaboost algorithm adaboost algorithm includes the use of very short (1-level) decision trees as weak learners added in a sequential manner to the set. every one of the consequent models tries correcting predictions that have been made by the model before it in a sequence. it combines several of the average or weak predictors for the purpose of building strong predictor [16]. h. performance measure performance measure of different machine learning algorithms is analyzed by considering measures such as [17]. • confusion matrix the confusion matrix is a table used in performance measures that helps in easy visualization as well as in distinguishing true positives, true negatives, false positives and false negatives. • accuracy accuracy measure is calculated by considering the ratio of the observations that have been correctly predicted to total number of the observations. accuracy = 𝑇𝑃+𝑇𝑁 𝑇𝑃+𝑇𝑁+𝐹𝑃+𝐹𝑁 (2) • precision it represents the percentage of true positives out of all the predictions. precision = 𝑇𝑃 𝑇𝑃+𝐹𝑃 (3) • sensitivity out of the total positive, what percentage are predicted positive. sensitivity = 𝑇𝑃 𝑇𝑃+𝐹𝑁 (4) • specificity – it represents true negative rate which is the proportion of the negative tuples which have been identified correctly. specificity = 𝑇𝑁 𝑇𝑁+𝐹𝑃 (5) iv. results and discussion on the implementation of algorithms that have been mentioned in previous section, the following results have been obtained: table1. confusion matrix actual / predicted normal abnormal normal tp fn abnormal fp tn table 2. confusion matrix of ann without rfe feature selection actual / predicted normal abnormal normal 15 3 abnormal 5 60 table 3. confusion matrix of ann with rfe feature selection actual / predicted normal abnormal normal 16 1 abnormal 5 61 table 4. performance measure of ann model without and with rfe feature selection model accuracy precision sensitivity specificity ann without 90.36% 75% 83.3% 92.3% vol. 4, no.2 july 2023 | 92 rfe ann with rfe 92.77% 76.1% 94.1% 92.4% table 5. confusion matrix of adaboost without rfe feature selection actual / predicted normal abnormal normal 12 5 abnormal 6 60 table 6. confusion matrix of adaboost with rfe feature selection actual / predicted normal abnormal normal 13 4 abnormal 5 61 table 7. performance measure of adaboost model without and with rfe feature selection model accuracy precision sensitivity specificity adaboost without rfe 86.74% 66.6% 70.5% 90.9% adaboost with rfe 89.15% 72.2% 76.4% 92.4% fig. 4. show accuracy of ann and adaboost models without and with rfe feature selection fig. 4. accuracy of ann and adaboost models v. conclusions this work presented a model for prediction of liver disease occurrence probability. the analyses and evaluations of suggested model have shown that it’s highly sufficient and easy to utilize and implement. two ml algorithms have been applied to ilpd data-set for classified liver patients. in the data preprocessing issue of imbalanced class distribution, an oversampling technique (random over sampling) is used, and used the z-score outlier detection technique to detect the outliers and handle those outliers using robust scaling. then applied rfe feature selection specifies the number of characteristics to be chosen is used for achieving better performance and for achieving an enhanced result, we have applied ann and adaboost algorithms. from the analysis of experimental results, the ann algorithm has achieved the highest accuracy of 92.77%. 0,82 0,84 0,86 0,88 0,9 0,92 0,94 ann without rfe ann with rfe adboost without rfe adboost with rfe accuracy models vol. 4, no.2 july 2023 | 93 references [1] h. hartatik, m. b. tamam, and a. setyanto, “prediction for diagnosing liver disease in patients using knn and naïve bayes algorithms,” 2020 2nd int. conf. cybern. intell. syst. icoris 2020, pp. 1–5, 2020, doi: 10.1109/icoris50180.2020.9320797. [2] m. a. kuzhippallil, c. joseph, and a. kannan, “comparative analysis of machine learning techniques for indian liver disease patients,” 2020 6th int. conf. adv. comput. commun. syst. icaccs 2020, pp. 778–782, 2020, doi: 10.1109/icaccs48705.2020.9074368. [3] m. a. khadija and n. a. setiawan, “detecting liver disease diagnosis by combining smote, information gain attribute evaluation, and ranker,” itsmart j. teknol. dan inf., vol. 9, no. 1, pp. 13–17, 2020. [4] a. s. jaddoa, z. tariq, and m. al-ta, “comparison of data mining algorithms for diagnosis of diabetes mellitus,” vol. 10, no. 2, pp. 1–8, 2021. [5] r. ahmed, s. jaddoa, p. ziyad, and t. mustafa, “diagnosis of diabetes mellitus using hybrid techniques for feature selection and classification,” pp. 1650–1663, 2021. [6] s. jain, r. sharma, and r. rajkamal, “easychair preprint classification of liver diseases using intelligent techniques classification of liver diseases using intelligent techniques,” 2021. [7] g. jamila, g. m. wajiga, y. m. malgwi, and a. h. maidabara, “a diagnostic model for the pediction of liver cirrhosis using machine learning teachniques,” comput. sci. it res. j., vol. 3, no. 1, pp. 36–51, 2022, doi: 10.51594/csitrj.v3i1.296. [8] g. s. harshpreet kaur, “the diagnosis of chronic liver disease using machine learning techniques,” inf. technol. ind., vol. 9, no. 2, pp. 554–564, 2021, doi: 10.17762/itii.v9i2.382. [9] m. ghosh et al., “a comparative analysis of machine learning algorithms to predict liver disease,” intell. autom. soft comput., vol. 30, no. 3, pp. 917–928, 2021, doi: 10.32604/iasc.2021.017989. [10] n. nahar, f. ara, m. a. i. neloy, v. barua, m. s. hossain, and k. andersson, “a comparative analysis of the ensemble method for liver disease prediction,” iciet 2019 2nd int. conf. innov. eng. technol., pp. 23–24, 2019, doi: 10.1109/iciet48527.2019.9290507. [11] s. afrin et al., “supervised machine learning based liver disease prediction approach with lasso feature selection,” bull. electr. eng. informatics, vol. 10, no. 6, pp. 3369–3376, 2021, doi: 10.11591/eei.v10i6.3242. [12] n. tanwar and k. f. rahman, “machine learning in liver disease diagnosis: current progress and future opportunities,” iop conf. ser. mater. sci. eng., vol. 1022, no. 1, 2021, doi: 10.1088/1757899x/1022/1/012029. [13] r. c. poonia et al., “intelligent diagnostic prediction and classification models for detection of kidney disease,” healthc., vol. 10, no. 2, 2022, doi: 10.3390/healthcare10020371. [14] s. kefelegn, “prediction and analysis of liver disorder diseases by using data mining technique: survey,” vol. 118, no. 9, pp. 765–770, 2017, [online]. available: http://www.ijpam.eu. [15] s. gupta, g. karanth, n. pentapati, and v. r. b. prasad, “a web based framework for liver disease diagnosis using combined machine learning models,” proc. int. conf. smart electron. commun. icosec 2020, no. icosec, pp. 421–428, 2020, doi: 10.1109/icosec49089.2020.9215454. [16] a. khatavkar, p. potpose, and p. pandey, “smart health prediction system,” vol. 5, no. 02, pp. 1550–1552, 2017. [17] b. k. mengiste, h. k. tripathy, and j. k. rout, “analysis and prediction of cardiovascular disease using machine learning techniques,” lect. notes electr. eng., vol. 708, no. 2, pp. 133– 141, 2021, doi: 10.1007/978-981-15-8685-9_13. . vol. 5, no.1 januari 2024 | 9 taking into account imprecision in the modeling voter in a multi-agent environment michel milambu belangany1*, pièrre kafunda katalayi2, eugène mbuyi mukendi3 1,2,3 department of mathematics and computer science, faculty of sciences and technology, university of kinshasa, 1,2,3 michel.milambu@unikin.ac.cd, pierre.kafunda@unikin.ac.cd, eugenembuyi@gmail.com abstract an electoral system is a set of individuals considered as agents in a multi-agent system in which voters communicate with each other and with the environment. in such a system, it is often difficult to understand the behavior of an agent that we call a voter. this is why, in this paper, we use fuzzy set theory as an approach to model the behavior of an imprecise voter in an electoral environment. it will be just a question of presenting a model of a voter with fuzzy behavior using mathematical approaches in this environment considered as a multi-agent environment and to propose the algorithms as the tools of computer modeling. keywords : voter, fuzzy voter, multi-agent system, electoral environment, modeling. i. introduction an electoral system is a complex system in which voters, candidates, and agents are involved and all of these people communicate. such a system is considered a multi-agent system. in such a system, it is often difficult to determine a voter's position or choice with respect to a candidate's vote. in boolean logic, an element belongs or does not belong to a given set: this is the "all or nothing". zadeh (1965) notes that most of the time, the objects encountered in the real world do not have precise criteria of belonging. he then tried to get out of this boolean logic by introducing the notion of weighted membership. he defines the fuzzy set as a class of objects with a continuum of degrees of membership to this class. such a set is characterized by a membership function that associates to each object a degree of membership between zero and one. in the case of a specific situation, each voter is assigned to exactly one candidate, which means that the degree of membership of the voter is 1 in the case of that candidate and 0 in all other situations. voter membership in candidates is thus mutually exclusive. on the other hand, a fuzzy voter allows belonging to several candidates at the same time; moreover, each voter has degrees of belonging that express how much this voter belongs to the different candidates. ii. related work multi-agent structures have been developed in the context of the management of complex systems such as electoral, epidemiological environments, in order to locate groups of agents, an agent in a group to monitor its behavior. these systems have allowed the understanding of areas in which there is communication, cooperation and collaboration between entities. in personality detection using contextbased emotions in cognitive agents. similarly, multi-agent systems used for search and rescue applications. researchers have difficulty converging on an unambiguous definition of notions such as interpretation or explanation, which are often (and wrongly) used interchangeably. moreover, despite the robust metaphors that multi-agent system (mas) could easily provide to address such a challenge, and agent-oriented perspective on the topic is still lacking. thus, this paper proposes an abstract and p-issn : 2715-2448 | e-issn : 2715-7199 vol.5 no.1 januari 2024 buana information technology and computer sciences (bit and cs) mailto:michel.milambu@unikin.ac.cd vol. 5, no.1 januari 2024 | 10 formal framework for xai-based mas, reconciling notions and results from the literature [6]. online social networks are known to lack adequate support for multi-user privacy. they present an agent architecture that aims to help users manage multi-user privacy conflicts. by considering the personal utility of content sharing and the individually preferred moral values of each user involved in the conflict, expri identifies the best collaborative solution by applying practical reasoning techniques. such techniques provide the agent with the cognitive process necessary for explicability [4]. in the race to automate, distributed systems are required to perform increasingly complex reasoning to cope with dynamic, often non-human controlled tasks. on the one hand, systems dealing with tight time constraints in safety-critical applications used to focus mainly on predictability, leaving little room for complex planning and decision making processes. indeed, real-time techniques are most effective in predetermined, constrained, and controlled scenarios [3]. theory of mind is generally defined as the ability to attribute mental states (e.g., beliefs, goals) to oneself and others. building and expanding on previous work by providing an account of the explanation in terms of agents' beliefs and the mechanism by which agents revise their beliefs given the possibility [11]. as autonomous agents become more autonomous, ubiquitous and sophisticated, it is vital that humans have effective interactions with them. therefore, these autonomous agents should be able to explain their behavior and decisions before humans can trust them. this paper focuses on analyzing human understanding of the behavior of explainable agents [1]. with the apparent societal need to design complex autonomous systems whose decisions and actions are humanly intelligible, the study of explainable artificial intelligence, and with it, research on explainable autonomous agents, a way to facilitate such studies by implementing explainable agents and multi-agent systems that (i) can be deployed as static files, not requiring server-side code execution, thus minimizing administrative and operational overhead, and (ii) can be integrated into web and other compatible user interfaces [22]. a network-oriented modeling approach for voting behavior in the 2016 us presidential election. a network-oriented computational model is presented for voting intentions over time, specifically for the race between donald trump and hillary clinton in the 2016 u.s. presidential election. emphasis was placed on the role of social and mass communication media and statements made by donald trump or hillary clinton during their speeches. the objective was to study the influence on voting intentions and the final vote. a sentiment analysis was conducted to test whether the statements were high or low language intensity [9]. the use of adaptive temporal-causal networks to model and simulate the development of mutually interacting opinion states and connections between individuals in social networks. the focus is on adaptive networks combining the homophily principle with the plus becomes plus principle. the model was used to analyze a dataset of opinions on alcohol and tobacco use and friendship [18]. in recent years, social networks have been increasingly used to study political opinion formation, monitor election campaigns, and predict election outcomes, as they are capable of generating a huge amount of data, usually in textual and unstructured form. the authors aim to collect and analyze data from twitter messages identifying emerging trends in topics related to a constitutional referendum that recently took place in italy in order to better understand and predict its outcome [17]. competitive multi-agent systems (mas) are inherently difficult to control due to agent autonomy and strategic behavior, which is especially the problem when there are system-level goals to achieve or specific environmental states to avoid. existing solutions for this task mainly assume specific knowledge about agents' preferences, utilities and strategies, neglecting the fact that actions are not always directly related to agents' true preferences, but may also reflect the anticipated behavior of competitors, be a concession to a superior adversary or simply be intended to deceive other agents. the authors propose a new approach to governance of competitive mas that relies exclusively on publicly observable actions and transitions, and uses the knowledge gained to deliberately restrict the action spaces, thereby achieving the system's goals while preserving a high level of autonomy for the agents [13]. vol. 5, no.1 januari 2024 | 11 in a new complex fuzzy inference system with fuzzy knowledge graph and extensions in decision making, the authors say that, complex fuzzy theory has strong practical implication in many real-world applications. complex fuzzy inference system (cfis) is a powerful technique to overcome the challenges of uncertain and periodic data [10]. in complex cubic fuzzy aggregation operators with applications in group decision making, the cubic set and complex fuzzy set are presented as two useful tools that have been successfully used to deal with fuzziness and uncertainties (xiaoqiang, yameng, zichang, feng & wu, 2020). in formalizing fuzzy control in possibility theory via rule extraction shows that possibility system has recently been recognized as a potential foundational theory for fuzzy theory. set theory although the concept of possibility is derived from the membership function of fuzzy sets. as an application of fuzzy set theory, fuzzy control has been widely used in engineering practices, where the control of laws is described by fuzzy if-then rules [20]. in real life, there will be many uncertainty problems, one of which is due to the vagueness of the concept of things, i.e., it is difficult to determine whether an object conforms to the concept. this situation largely exists in some states, phenomena, parameters and interrelationships between things. for such uncertain events with heavy subjective influencing factors and incomplete data, fuzzy methods should be used to cope with them [11]. none of the aforementioned works have used fuzzy sets in the electoral domain. in this paper, an agent is a voter, a candidate in an electoral environment capable of speaking or expressing an opinion with different agents on an occasion within the environment and perceiving its environment, manipulating the objects in the environment including the election kits. referring to the expression of opinions, it can be said that some voters have an imprecise opinion. these kinds of voters are fuzzy voters who are even the subject of this article in which we model the behavior of a fuzzy voter based on his imprecise language. iii. methods 1. material in this article, we used anaconda navigator as a utility particularly jupyter to manipulate the libraries of the python language in particular matplotlib which is a python library which allowed us to visualize the data and to draw the curves [16]. numpy is a python library that contains functions related to the manipulation of our data which are represented in matrix form (two dimensional arrays, vectors)[16]. etc..., the draw.io environment to represent the electoral system composed of individuals (voters, candidates and agents ...). the notepad that contains the data used and the ms excel that allowed us to organize the data and compare the curves. 2. methods we use the theory of fuzzy subsets which will allow us to present the imprecise behavior of a voter in an electoral system. let x be a reference set and let x be any element of x. a fuzzy set a of x is defined as the set of couples : 𝐴 = {(𝑥, 𝜇𝐴(𝑥)), 𝑥 ∈ x} (1) where : 𝜇𝐴: 𝑋 → [0, 1] (2) thus, a fuzzy set a of x is characterized by a membership function that associates, to each element x of x a real in the interval [0, 1]; 𝜇𝐴(𝑥) represents the degree of membership of x to a. thus, the closer the value of 𝜇𝐴(𝑥) is to unity, the higher the degree of membership of x to a [2]. if we have: 𝜇𝐴: 𝑋 → {0, 1 } we find the boolean case: vol. 5, no.1 januari 2024 | 12 either x belongs to 𝐴(𝜇𝐴 = 1) or it does not belong to 𝐴(𝜇𝐴 = 0). and the following case is very useful in the sense that an element belongs partially: let x belong partially to 𝐴(0 < 𝜇𝐴(𝑥) < 1) it is important to specify that the fuzzy set is considered as empty if the membership degrees of all the elements of the universe are all equal to zero. 𝐴 = ∅ ⇔ 𝜇𝐴 (𝑥) = 0, ∀𝑥 ∈ 𝑋 (3) two fuzzy sets are equal if their membership degrees are equal for all elements of the reference set, i.e., if both fuzzy sets have the same membership function [5]. two fuzzy sets a and b, defined on the same reference set x are equal if: a = b ⇔ 𝜇𝐴(𝑥) = 𝜇𝐵(𝑥), ∀𝑥 ∈ 𝑿 (4) a. modeling this complex system in figure 1 we call the electoral system which is composed of individuals who communicate. to model a fuzzy voter, we use the theory of fuzzy subsets as below: figure 1. complex system composed of the voters. thus, referring to the theory of fuzzy subsets above, our model can be represented as follows: μ𝐸𝑓 (𝑒) ∶ 𝑆𝑒 → [0, 1] (5) where : 𝑆𝑒:universe of discourse or reference set e ∶ voter (any element of 𝑆𝑒) e𝑓 ∶fuzzy subset of 𝑆𝑒 μ𝐸𝑓 (𝑒) ∶membership function that measures the degree to which e belongs to 𝐸𝑓 we then define this membership of a voter by : 𝜇𝐸𝑓 (𝑒) = 1 if e belongs completely to 𝐸𝑓 𝜇𝐸𝑓 (𝑒) = 0 if e does not belong to 𝐸𝑓 0 < 𝜇𝐸𝑓 (𝑒) < 1 if e belongs partially to 𝐸𝑓 taking the array of fuzzy values 𝑇𝑣𝑓 of size n from which each part can be removed as a vector of the fuzzy values of a voter below: vol. 5, no.1 januari 2024 | 13 table 1. matrix of fuzzy membership values for voters 𝒆𝟏 0.0000 0.01 0.04 0.02 … 0.3 𝒆𝟐 0.1 0.006 0.03 0.05 … 0.005 𝒆𝟑 1 1 1 1 … 1 ⋮ ⋮ ⋮ ⋮ ⋮ ⋮ ⋮ 𝒆𝒎 0.007 0.07 0.7 0.8 … 0.9 b. proposal of the algorithms proposal algorithm input 𝑆𝑒 = {𝑒1,𝑒2, … , 𝑒𝑛}; 𝑄𝑓 = [0, 1]; 𝐸𝑓 begin 1. repeat 2. for i = 1 to n do 3. read : {𝑆𝑒} 4. the fuzzy quantifier 𝑄𝑓 fuzzyfie the {𝑒𝑖} 5. if 0 < 𝜇𝐸𝑓 (𝑒𝑖) < 1 then 6. 𝑒𝑖 𝑎𝑑𝑚𝑖𝑡𝑠 𝑠𝑒𝑣𝑒𝑟𝑎𝑙 𝑐𝑎𝑛𝑑𝑖𝑑𝑎𝑡𝑒𝑠 7. 𝑒𝑖 𝑖𝑠 𝑓𝑢𝑧𝑧𝑦 8. endif 9. endfor 10. end repeat 11. end. the second proposed algorithm on the fuzzy behavior of a voter is the following: fuzzy voter algorithm input 𝑇𝑣𝑓 = {𝑣𝑓1,𝑣𝑓2, … , 𝑣𝑓𝑛} , 𝑆𝑒 = {𝑒1,𝑒2, … , 𝑒𝑚}, 𝐸𝑓 begin 1. repeat 2. for i = 1 to n do 3. if 𝑇𝑣𝑓[𝑖] 𝑖𝑠 𝑑𝑖𝑓𝑓𝑒𝑟𝑒𝑛𝑡 𝑡𝑜 𝑇𝑣𝑓[𝑖 + 1] then 4. 𝑒𝑖 𝑖𝑠 𝑓𝑢𝑧𝑧𝑦 5. 𝑓𝐸𝑓 (𝑒𝑖) ≠ 𝑓𝐸𝑓 (𝑒𝑖+1) : so it is fuzzy 6. else 7. 𝑓𝐸𝑓 (𝑒𝑖) == 𝑓𝐸𝑓 (𝑒𝑖+1) : is not fuzzy, ∀𝑒𝑖 ∈ 𝑆𝑒 8. endif 9. endrepeat 10. end iv. result and discussion the imprecision zone is the part covered by the fuzzy voters, while the precision zone contains the voters with a precise choice. vol. 5, no.1 januari 2024 | 14 figure 2. distribution of voters the representation of the fuzzy part of voters in the figure below : figure 3. the fuzzy voters fuzzy behavior of voters represented by two curves of which in blue the language of the voter is totally fuzzy and in red the language of the voter is partially fuzzy figure 4. the fuzzy voter membership function 𝜇𝑒𝑖 (𝐿𝑓) < 50 ∶ 𝑡ℎ𝑒𝑛 𝑡ℎ𝑒 𝑣𝑜𝑡𝑒𝑟 𝑖𝑠 𝑓𝑢𝑧𝑧𝑦 𝜇𝑒𝑖 (𝐿𝑓) = 50 ∶ 𝑇ℎ𝑒𝑛 𝑡ℎ𝑒 𝑣𝑜𝑡𝑒𝑟 𝑑𝑜𝑒𝑠 𝑛𝑜𝑡 𝑏𝑒𝑙𝑜𝑛𝑔 50 < 𝜇𝑒𝑖 (𝐿𝑓) ≤ 100 ∶ 𝑇ℎ𝑒𝑛 𝑡ℎ𝑒 𝑣𝑜𝑡𝑒𝑟 𝑡𝑜𝑡𝑎𝑙𝑙𝑦 𝑏𝑒𝑙𝑜𝑛𝑔𝑠 vol. 5, no.1 januari 2024 | 15 𝐼𝑓 𝐿𝑓 = 50 ∶ 𝜇𝑒𝑖 (𝐿𝑓) = 0 𝐼𝑓 𝐿𝑓 < 50 ≤ 𝜇𝑒𝑖 (𝐿𝑓) ∶ 𝜇𝑒𝑖 (𝐿𝑓) − 𝐿𝑓 𝜇𝑒𝑖 (𝐿𝑓) − 50 𝐼𝑓 𝐿𝑓 ≥ 60 ∶ 𝜇𝑒𝑖 (𝐿𝑓) = 1 where : 𝑳𝒇: 𝑓𝑢𝑧𝑧𝑦 𝑙𝑎𝑛𝑔𝑎𝑔𝑢𝑒 figure 5. behavior of a fuzzy voter over the course of a day figure 5 above, presents the evolution of the membership behavior of a fuzzy voter by day, i.e. each day he has a position for each candidate. in fact, looking at this figure, at the beginning this voter does not accept any candidate, then he tries to give a position for the first candidate at 0,01 and the following day 0,02 for another one the third day he gives less than 0,01 to the starting candidate therefore each day his position changes at the thirty second day he gives a position with a total membership of 1 and after he changes his value again therefore he does not have a total membership. figure 6. behavior of a fuzzy voter over time in this figure 6, the fuzzy agent or voter shows different positions every 5 days interval. as we can see in this figure 6 where it rises with a position of 0.2 for a candidate and after the next few days it falls to less than 0.2 for another candidate. so for each interval of 5 days there is a position of this voter that is different from the previous one that makes him a fuzzy voter. vol. 5, no.1 januari 2024 | 16 figure 7. behavior of a fuzzy voter over the course of a month in this figure 7, the voter fumbles over an interval of months taking different positions on the different candidates. here, a voter blurs in the first three months his position is below 0.2 for the candidates and the last two months in the interval of 5 months he is below 0.1 and in the next interval he belongs totally to one candidate with 1 in the interval of 10 to 15 months, he makes 0.9. so for each interval of months it changes positions several times. and finally of account in interval of months it changes values of measure of degree of confidence to the candidates. we will notice that, this voter is inconstant. therefore, he is a fuzzy voter whose behavior is imprecise. figure 8. behavior of a fuzzy voter and a normal voter in this figure 8, we see the behavior of a fuzzy voter and a normal or accurate voter. looking at this figure, the fuzzy voter seems to be unstable or imprecise because it changes its behavior over time so for each sequence of weeks it displays some behavior. on the other hand, a precise or normal voter remains constant by keeping a certain position during all the weeks. for this purpose, it is necessary to identify the electoral environment in which a voter evolves after having noted among many voters that not all of them are stable, i.e. have a well defined position. v. conclusion this article focused on the consideration of the imprecision in the modeling of a voter in order to be able to follow his behavior throughout an electoral process. it was in fact a question of being able to use the theory of fuzzy subsets in order to present a model capable of defining or showing the evolution of the behavior of a voter with respect to the candidates in an electoral system. the use of imprecision in this modeling showed that during an electoral process a voter can be imprecise in the sense that every day, every week or every month he has a thought about the candidates with a certain value of his membership function which defines the degree of confidence of this voter with regard to the candidates. this approach offers significant advantages using our algorithms proposed in this article to assure election candidates that not everyone who follows you is necessarily your voter because as he follows you so he follows another candidate. so for each candidate, a fuzzy voter has a confidence measure value that we call the candidate membership function of a fuzzy voter that can change by day, week or month so that on the day of the vote the voter can go without a specific choice of candidate. vol. 5, no.1 januari 2024 | 17 references [1] avleen, m., samanta, k., & kary, f. (2020). explainable agents for less bias in human-agent decision making. springer nature switzerland pp. 129–146 https://doi.org/10.1007/978-3-03051924-7_8 [2] didier, d., & henri, p. fuzzy sets and systems : theory and applications. mathematics in science and engineering volume 144. cnrs, languages and computer systems (lsi) université paul sabatier toulouse, france [3] francesco, a., paolo, g., amro, n., michael, i., schumacher, & davide, c. (2020). in-time explainability in multi-agent systems: challenges, opportunities, and roadmap. springer pp. 39–53 https://doi.org/10.1007/978-3-030-51924-7_3 [4] francesca, m., s,tefan, s., jose, m. such, & peter, m. (2020). agent expri : licence to explain. springer https://doi.org/10.1007/978-3-030-51924-7_2 [5] george, j., klir, & bo, y. (2013). fuzzy sets and fuzzy logic : theory and applications. printed in the united states of america 10 9 8 7 6 5 4 3 2 1 isbn 0-13-101171-5 corpsales@prenhail.com [6] giovanni, c., michael, i., schumacher, andrea, o., & davide, c. (2020). agent-based explanations in ai: towards an abstract framework. springer nature switzerland https://doi.org/10.1007/978-3-030-51924-7_1 [7] janusz kacprzyk & baoding liu, studies in fuzziness and soft computing, volume 239, springerpolish academy of sciences, 2nd edition. issn 1434-9922 doi 10.1007/978-3-54089484-1 [8] lindsay, s., & julie, a. (2020). a situation awareness-based framework for design and evaluation of explainable ai. springer nature switzerland https://doi.org/10.1007/978-3-03051924-7_6 [9] linford, g., jan, t., & roos, v. (2018). a network-oriented modeling approach to voting behavior during the 2016 us presidential election. © springer international publishing. advances in intelligent systems and computing 619, doi 10.1007/978-3-319-61578-3_1 [10] luong thi hong, l., tran manh, t., tran thi, n., le hoang, s., nguyen long, g., vo truong nhu, n., & pham van, h. (2020). a new complex fuzzy inference system with fuzzy knowledge graph and extensions in decision making. ieeeaccess doi :10.1109/access.2020.3021097 [11] maayan, s., toryn, q., klassen, a., sheila, a., & mcilraith. (2020). towards the role of theory of mind in explanation. springer nature switzerland https://doi.org/10.1007/978-3-030-519247_5 [12] manuel h., marco, pérez-h., ajith kumar, p., & joaquín, i. (2020). a review on control and optimisation of multi-agent systems and complex networks for systems engineering. preprints.org doi:10.20944/preprints202001.0282.v1 [13] michael, p., christian, b., & heiner, s. (2021). governing black-box agents in competitive multi-agent systems. springer nature switzerland pp. 19–36 https://doi.org/10.1007/978-3030-82254-5_2 [14] muhammad, g., m. haris, m., dilshad, a., & nasreen, k. (2020). a novel applications of complex intuitionistic fuzzy sets in group theory. ieeeaccess doi : 10.1109/access.2020.3034626 [15] nafiseh, j., ali, h., hossein, & sayyadi, t. (2021). designing an intuitionistic fuzzy network data envelopment analysis model for efficiency evaluation of decision-making units with two-stage structures. hindawi advances in fuzzy systems, article id 8860634, https://doi.org/10.1155/2021/8860634 [16] patrick, f., & pierre, p. (2020). cours de python. université de paris, france https://python.sdv.univ-paris-diderot.fr/ https://doi.org/10.1007/978-3-030-51924-7_1 vol. 5, no.1 januari 2024 | 18 [17] shira, f., & debora, s. (2018). using twitter data to monitor political campaigns and predict election results. springer international publishing advances in intelligent systems and computing 619, doi 10.1007/978-3-319-61578-3 19 [18] sven van den, b., simon, h., goos, & jan t. (2018). understanding homophily and morebecomes-more through adaptive temporal-causal network models. © springer international publishing advances in intelligent systems and computing 619, doi 10.1007/978-3-31961578-3_2 [19] tomoki, y., yuki, m., & toshiharu, s. (2021). path and action planning in non-uniform environments for multi-agent pickup and delivery tasks. springer nature switzerland pp. 37– 54 https://doi.org/10.1007/978-3-030-82254-5_3 [20] wei, m. (2020). formalization of fuzzy control in possibility theory via rule extraction. ieeeaccess digital object identifier 10.1109/access.2019.2928137 [21] xiaoqiang, z., yameng, d., zichang, h., feng, y., & wu l. (2020). complex cubic fuzzy aggregation operators with applications in group decision-making. ieeeaccess doi :10.1109/access.2020.3044456 [22] yazan, m., timotheus, k., igor, h., tchappi, amro, n., stéphane, g., & christophe n. (2020). explainable agents as static web pages: uav simulation example. springer nature switzerland pp. 149–154 https://doi.org/10.1007/978-3-030-51924-7_9 vol. 4, no.2 july 2023 | 54 using i-hubs for bridging the gap of digital divide in rural kenya samuel w lusweti 1, kelvin k omieno2 1 department of information technology, school of computing and informatics, masinde muliro university of science and technology 2 department of information technology, school of computing and information technology, kaimosi friends university. 1lusweti015@gmail.com, 2komieno@kafuco.ac.ke abstract the world is moving towards digital economy where almost everything being done today is digitally controlled because necessity is the mother of innovation. everybody is striving to attain digital stability as a lot of revenue is generated in the digital world. digital divide therefore becomes so disadvantageous to people left without access to computers and the internet. in this paper, researchers discuss the role of kenyan innovation hubs in closing the gap between those who have access to the internet and computers and those who do not. the paper discuss the world bank projection of the gdp emanating from the use of icts and the challenges facing innovation. government support plays a key role in ensuring that the people secluded from icts are able to access these services especially those in rural areas. this research found that in kenya, innovation hubs have helped the citizens staying in rural areas to gain access to internet and develop their ideas and innovations as well as undergo mentorship. nonetheless, a lot of support is needed from the kenyan government through the launching of more innovation hubs especially in rural areas that can help improve the online business, innovation and thus increase the gdp from icts. key words: i-hubs, digital divide, innovation, gdp, internet abstrak dunia bergerak menuju ekonomi digital di mana hampir semua yang dilakukan saat ini dikendalikan secara digital karena kebutuhan adalah ibu dari inovasi. semua orang berusaha untuk mencapai stabilitas digital karena banyak pendapatan dihasilkan di dunia digital. oleh karena itu, kesenjangan digital menjadi sangat tidak menguntungkan bagi orang-orang yang dibiarkan tanpa akses ke komputer dan internet. dalam makalah ini, peneliti membahas peran hub inovasi kenya dalam menutup kesenjangan antara mereka yang memiliki akses ke internet dan komputer dan mereka yang tidak. makalah ini membahas proyeksi bank dunia terhadap pdb yang berasal dari penggunaan tik dan tantangan yang dihadapi inovasi. dukungan pemerintah memainkan peran kunci dalam memastikan bahwa orang-orang yang terpisah dari tik dapat mengakses layanan ini terutama di daerah pedesaan. penelitian ini menemukan bahwa di kenya, pusat inovasi telah membantu warga yang tinggal di daerah pedesaan untuk mendapatkan akses ke internet dan mengembangkan ide dan inovasi mereka serta menjalani bimbingan. meskipun demikian, banyak dukungan yang dibutuhkan dari pemerintah kenya melalui peluncuran lebih banyak hub inovasi terutama di daerah pedesaan yang dapat membantu meningkatkan bisnis online, inovasi, dan dengan demikian meningkatkan pdb dari tik. kata kunci: i-hub, kesenjangan digital, inovasi, pdb, internet p-issn : 2715-2448 | e-issn : 2715-7199 vol.4 no.2 july 2023 buana information technology and computer sciences (bit and cs) mailto:lusweti015@gmail.com mailto:komieno@kafuco.ac.ke vol. 4, no.2 july 2023 | 55 i. introduction the united nation’s world summit of information society (wsis), geneva 2003, tunis commitment, tunis 2005 and the copenhagen declaration agreed that ict is a key player in eradication of poverty and unemployment as it helps in creation of development-oriented information society [25]. the best way to make use of ict services is through research and innovation; the later becoming the most widely used method of technological advancement. innovation is the change in traditional practice for the remarkable success of humanity towards colonizing the earth, technology and product diversification [1] . the importance of innovation has thus made human beings to be cleverer and daring to make their lives more comfortable especially when it comes to technological advancements in various fields. ict has not been left out in this, as many countries around the world have resorted to innovation in ict. kenya has for almost 20 years committed itself to gain economic stability and global recognition through technological advancement. the commitment of kenya in technological developments started with vision 2030 economic blueprint [2]. in trying to meet the objectives of this economic blueprint, kenya has made a great milestone in technological advancement in the area of ict for years now. during the period of covid-19 pandemic, innovation was heightened in the country to keep its economy afloat amid the ravaging effects of the pandemic. the innovations which were majorly concentrated on use of ict for teaching and learning, making of local computer-controlled ventilators and online business transactions helped in fast-tracking the achievement of kenya’s ict policy vision 2030 [3]. innovation hubs are also an important component of technological advancement that are used to help countries attain economic stability and for interconnecting people in different parts of the world to interact and share knowledge and resources. i-hubs form a basis were innovation ideas are bred, incubated and hatched into solution centered systems basing on prevailing conditions in a given environment. a. barriers to innovation in developing countries developing countries are generally known to have low industrial development and with lower human development index as compared to other countries in the world. this means that the countries are still growing in terms of technological advancements. the united nations classifies countries according to their economic stability and growth as either developed or developing. kenya is classified under the category of developing countries among many other countries in africa, asia and latin america [4]. developing countries generally face a number of challenges that affect the broad area of development including agriculture, industry, education medicine and technology. the following are some of the challenges faced by developing countries derailing their industrialization process: 1) inadequate funding, 2) poor infrastructure, 3) insufficient support from governments, 4) lack of motivation for innovations, 5) inadequate research facilities, 6) lack of opportunities and poor cooperation among business firms [5]. innovation helps countries to be economically stable hence most countries classified as developed do well in innovation, creativity and invention. in kenya for example, innovations among its citizens in the early years of 21st century lagged behind due to inadequate funding, insufficient support from the government, volatility of innovation and inadequate research facilities. due to the volatile nature of technological innovations and the changing technological environments, it has been a challenge for kenyan innovators to strike the balance of mechanisms and expectations of the society on technological innovation [6]. this is especially difficult when working alone without government support and without i-hubs where people meet to bring onboard different ideas. however, during the wake of the novel corona virus disease, the country witnessed a lot of innovations and inventions in various areas including medicine, and ict [2]. b. digital divide in 1990s, digital divide was defined as the gap between those who have access to computers and the internet and those who do not have the access [7]. digital divide existed long before 1990s but the term was not being used before, instead information inequality or knowledge gap were the common terms used. the origin of this term is vol. 4, no.2 july 2023 | 56 connected to the us department of commerce national telecommunications and information administration (cntia) [8] which advices the us president on the telecommunication and information policies. digital divide therefore excludes some groups of people from the uptake of information communication technologies. according to [9], the following groups of people are excluded in icts by digital divide: low-income earners, elderly people, unemployed, people with low literacy level, rural and urban areas, single parents, people leaving with disabilities, women and children. in kenya, there exists digital divide between rural and urban areas, and between kenya (developing country) and developed countries. poor power supply in rural areas in kenya which predicates the internet and computers among other aspects make the internet to be less cost-effective [10] leading to poor distribution of the internet in these areas. this inadequacy of power, computers and the internet has made rural areas in kenya to lag behind in technology and innovation. figure 1 below shows the penetration and uptake of internet in kenya for a period of 30 years since 1990 to 2020. figure 1 sourced from world bank: individuals using internet (% of kenya’s population) [11]. from figure 1 above, there has been an exponential increase in the number of people in kenya accessing the internet for a period of 30 years. however, for that period a maximum of only 30% of kenyans are able to access the internet. this could be as a result of the large population of people leaving in rural areas who are unable to access the internet services. according to the un specialized agency for icts, the world is increasingly becoming digitized, and the economy of digital exclusion is more expensive than closing the gap of digital divide [12]. the urgency further stated that 60% of the global gdp is expected to be contributed by digital technologies by 2022 thus people who are unable to connect to digital systems highly risk being left trailing in economic development. figure 2 below sheds more light why ict has been important in kenya’s economy over time from 2013 to 2019. vol. 4, no.2 july 2023 | 57 figure 2 source: kenya national bureau of statistics figure 2 above shows that ict has been contributing to gross domestic product (gdp) growth of the country by 10.8% on average annually since the year 2016 thus proved to be a dynamic source of economic growth and job creation [13]. although there is a promising growth of gdp from ict in kenya going into the future, a lot has to be done in order to increase the input of ict on the country’s gdp so as to meet the projection of united nations specialized agency for icts of 60% although the country is already behind the schedule. the main cause of this derail in many countries is the gap between those who have an access to internet and computers and those who do not have any access. this is because by january 2022, only 42.0% of kenyans were able to use internet which has not reached in many rural areas as 71% of the population stay in rural areas [14]. to narrow this gap, kenya as a country has among many other practices invested in i-hubs that would then help fast track the contribution of icts towards the 60% ict gdp projection by the un. ii. methods this research paper adopted qualitative research method in which secondary data which is published online with information on innovation was used. thematic analysis methods were used to dissect, understand and analyze collected data in related areas. iii. results and discussions a. innovation hubs (i-hubs) i-hubs have also been a source of bringing people from different backgrounds together for exchange of ideas thereby enhancing co-working which has been applied in various fields around the world [15]. co-working spaces created by establishment of innovation hubs help knowledge workers to gather together, create knowledge and benefit from it by working alone yet staying together [16]. these hubs have consequently helped the freelancers to vol. 4, no.2 july 2023 | 58 come together for support on the art of freelancing and social connection. innovation hubs have also been used in many developed and developing countries to help enhance their technological advancements. in uk for example, the london innovation hub (lih) which started in the year 2005, has transformed into 80 hubs with membership of over 13,000 worldwide [17]. lih is especially important because everybody in the world today looks at london as a model of innovation, social entrepreneurship and development. consequently, in zambia (africa), the lusaka innovation founded in 2011 was the only hub in the country which was aimed at helping young graduates of computer science gain skills in programming [17]. the hub has evolved to become a centre for innovators to connect and collaborate in working on various viable business project models. in kenya, a number of innovation hubs have been started by various stake holders both in public and private sectors making kenya to become a tech hub in the entire of east africa [2]. the following are some of the renowned i-hubs started in kenya that are harnessing technical knowledge and ict innovation from kenyan youths to enhance technological advancements up to rural areas of the country. b. kenya: the silicon savannah silicon savannah is widely known as the hub for technological ecosystem making kenya a recognized it hub in east africa. kenya coined the title silicon savannah which means that the country provides a fruitful environment for government support for innovation through technology. the moniker term silicon savannah which springs from two ecosystems of california’s silicon valley is a center for high tech and innovation and grassland savanna which is an ecosystem characterized by tree and grasses forming a conducive environment for growth. the mobile technology and digital payments made a great step in earning kenya this name [18]. silicon savannah is home to one of the fastest internet speed in the world; thanks to the fiber optic cable laid undersea making big firms such as microsoft, intel, facebook and ibm an environment to invest in high tech technologies [19]. through silicon savannah, kenya is becoming a major technological hub in africa. in the year 2007, kenya’s iconic telecommunication company safaricom launched m-pesa which is a mobile money transfer service that earned kenya a global recognition for becoming the first banking service for mobile phones to be developed and used in developing countries [20].the mobile phone money transfer service has been in operation till today in kenya and beyond whereby in the year 2022 alone, there were over 19 billion m-pesa transactions in kenya [20]. after development of m-pesa app, later in the same year 2007, turbulent political instability in kenya led to the launch of ushahidi (witness) app which helped citizens to track the real time post-election violence outbreaks and report to police immediately [21]. silicon savana moniker coined by kenya, proves that the country is alive to technological advancement and has considerably invested in innovations. c. afrilabs and association of countrywide innovation hubs (acih) association of countrywide innovation hubs (acih) based outside of nairobi-kenya, has the objective of promoting activities of its member hubs and to support them build sustainable businesses in rural and semi-urban areas in kenya. afrilabs which is africa’s largest network of innovation hubs was started in the year 2011 with the aim of supporting african entrepreneurs, developers and innovators. it achieves this objective by proving them with a friendly co-working space and training them in readiness for developing and implementing innovative solutions facing the continent [22]. afrilabs and the (acih) signed a long-term partnership to co-support the rural innovations by creating training and mentorship programs to help grass root i-hubs in africa thrive [23]. “through this collaboration, we are looking to an inclusive innovation ecosystem in africa and becoming a source of prosperity for all by strengthening grass root hubs to foster rural and peri-urban innovation.” said acih chairperson. afrilabs supports over 340 innovation hubs across all the 52 countries in africa and empowers innovators and developers through technological trainings, legal and financial support. on the other hand, acih is a network of i-hubs in kenya found outside nairobi city with membership of over 56 technology and innovation hubs with the main objective of promoting activities of member hubs. acih does this by supporting their vision of building sustainable businesses in rural and second-tier towns in kenya [23] through expert opinion, joint research and engaging the government in policy frameworks on behalf of startups. vol. 4, no.2 july 2023 | 59 d. konza technopolis kenya’s vision 2030 economic blueprint paved way for the development of konza technology city project aiming at developing a technology innovation hub in africa [24]. the city is located in a rural area 60 kilometers south of nairobi capital city on the way to mombasa city and is still under construction. it is estimated to cost about 1.2 trillion kenyan shillings (us $14.5 billion) upon completion. according to konza city website, [25] it will be a world class city powered by the ict sector, superior infrastructure and governance systems that are business friendly. the city will sit on a 5,000-acre plot surrounding three rural counties; machakos, makueni and kajiado. the project will attract software developers, disaster recovery centers, data centers and will be home to a hi-tech university focusing on research and technology, tvet institutions, schools and hospitals and stadiums [25]. these investments will make konza a smart city and a hub of innovation bringing light to former rural areas. the konza technopolis development authority (kotda) developed the vision to make the city a “global technology and innovation hub” and the mission to “develop a sustainable smart city and an innovation ecosystem, contributing to kenya’s knowledge-based economy” [25]. with these objectives, the city will create 200,000 jobs by the end of the year 2030. many of the staff who will work in the city during construction and setup of ict infrastructure will be locals from the rural areas of the three counties making them interact with digital equipment and infrastructure. e. national government innovation hubs the national government of kenya through the ministry of ict in the year 2019 established 300 constituency innovation hubs across all the constituencies in the country [26]. this establishment was meant to help train the youths across 290 constituencies on online jobs through the use of ajira digital platform. this platform was launched by the same ministry of ict in the year 2018 to help over one million youths to work online annually. according to the principal secretary of the ministry of ict, the world is moving away from traditional physical jobs to multiemployer online jobs [26] . the innovation hubs are also meant to help the youth enhance their innovation capability which is in line with the country’s ict policy. the ministry was set to partner with constituency development fund boards to provide laptops and internet to youths from various constituencies living in both rural and urban areas [26]. through the use of these innovation hubs. the government of kenya was able to train over 150,000 youths on ajira digital platform by december 2022 increasing the number of youths working online to 1.9 million up from 1.2 million in the year 2021 [27]. due to establishment of the 300 innovation hubs in the country with at least once i-hub per constituency, many youths in rural kenya are also able to access computers and internet. this is because more than half of the constituencies in the country are in rural regions. f. county governments i-hubs kenya is divided into 47 county governments which run devolved central government programs. bungoma county is one of the rural county governments available in kenya. being a rural county, technological uptake and advancement is very low as compared to urban counties like nairobi. this is because there is poor distribution of electrical power, computers and the internet [10] bringing about undesired gap of digital divide between the two counties. in the quest to bridge this gap, the county government of bungoma in the year 2015 launched a us$2 million matili technology hub (mthub) during the ict convention and innovation forum [28]. the aim of the hub was to help tap innovative minds for the economic prosperity by use of ict in the county. “there is great emphasis across major cities to adopt technology solutions, but this time we are looking to have citizens at county level innovate and benefit from solutions that they themselves engineered” said the governor of bungoma county at that time. prior to the launch, the county government officials partnered with international universities to help set up the i-hub. they also had travelled to lagos nigeria with young tech entrepreneurs to participate in the demo africa event [28]. the event equips young techies with knowledge and skills relevant in setting up technology parks and innovation hubs as well us networking. in baringo county, there exists rift valley innovation centre (rvic) which is one of the newly established rural based ict innovation hubs [29]. it consists of ultra-modern ict center vol. 4, no.2 july 2023 | 60 with computers, internet and servers that can support over 200 users concurrently. rvic is an incubation center for young techies to add value to the community through entrepreneurship and business mentorship. iv. conclusion this paper presented importance of ict to countries in the quest to attain industrial and economic stability. the paper has discussed how innovation in ict can help fast-track the implementation of ict in various areas. however, the challenges facing ict penetration in kenya have been discussed. nevertheless, the paper as also how various stake holders are working to see that innovation is open to all amid these challenges. in the past, many innovators failed to achieve their dreams because they stayed in rural areas where there was poor internet connectivity, very few computers and lacked support and training. this problem brought about a disparity between those who had access to computers and the internet and those who did not have access (especially those leaving in rural areas). the disparity which is widely referred to as digital divide is undesirable as it makes the people to economically lag behind in the current information society. at the onset of digital innovations in kenya, only people leaving in urban areas were able to learn about using ict for coming up with innovations as well as showcase their innovations. however, in the recent past, the idea of innovation hubs and incubation hubs have greatly improved on how to harness the skills of young innovators regardless of where they stay. from the data collected in this paper, researchers discussed how innovation hubs have been used to help kenyan innovators staying in rural areas to have access to computers and the internet. this access to internet has become a ladder towards bridging the gap of digital divide that has for long existed between the rural and urban areas. although, data from the world bank and kenya bureau of statistics show that even though the population of kenyans who can access internet is growing, their percentage is still too low and the number of innovation hubs is very low. more and more innovation hubs need to be put into villages by the kenyan government and other stake holders to help many kenyans who cannot access or afford the internet gain free access so that they can achieve their innovation dreams through ict. references [1] l.k racheal. et al., "eureka!: what is innovation, how does it develop, and who does it?," child development, vol. 87, no. 5, p. 1505–1519, october 2016. [2] garage., “the rise of global tech hubs in kenya// the impact of startup ecosystem” [online]. available: https://nairobigarage.com/the-rise-of-global-tech-hubs-in-kenya/ [accessed 21 june 2022]. may 2022. [online]. [3] s. w. lusweti and c.o. odoyo., "covid-19 pandemic as an accelarator toward attainment of ict policy-kenya vision 2030," computer science information technology, vol. 10, no. 3, pp. 30-35, october 2022. [4] un, “country classifications”, world economic situation and prospects 2022 . [accessed 20 july 2022]. available: https://www.un.org/development/desa/dpad/wpcontent/uploads/sites/45/wesp2022_annex.pdf [5] l.n moshood and o.f dotun, "barrier to innovation in develeping countries' firms: evidence from nigerian small and medium scale enterprises," european scientific journal, vol. 11, no. 19, pp. 1857-7881, july 2015. [6] n.g hottensiah, innovation challenges encountered by smalland medium enterprises in nairobi, nairobi: university of nairobi, 2017. https://www.un.org/development/desa/dpad/wp-content/uploads/sites/45/wesp2022_annex.pdf https://www.un.org/development/desa/dpad/wp-content/uploads/sites/45/wesp2022_annex.pdf vol. 4, no.2 july 2023 | 61 [7] a.g.m jan, "digital divide research, achievements and shortcomings," universityof twente, department of communication, , vol. 34, pp. 221-235, 2006. [8] n.t.i.a, “falling through the net: defining the digital divide”, 1999. [online]. available: http://www.ntia.doc.gov/ntiahome/fttn99/contents.html. [9] c. rowna., "addressing the digital divide," mcb university press, vol. 25, no. 5, pp. 311-320, 2001. [10] s. mutula, "internet connectivity and services in kenya: current developments," mcb up limited, vol. 20, no. 6; doi 10.1108/02640470210453949, pp. 466-472, 2002. [11] w. bank.,“individuals using the internet(% population)-kenya” , 2020. [online]. available: https://data.worldbank.org/indicator.it.net.user.zs?end=2020&locations=ke&start=1990&view=chart. [accessed 21 november 2022]. [12] unitu, november 2021. “bridging the digital divide with innovative finance and business models.”, november 2021. [online]. available: https://www.itu.int/hub/2021/11/bridging-the-digital-divide-with-innovative-finance-and-business-models/ [13] w. bank., "kenya economic update," policies to support kenya's digital transformation, vol. 20, 2019. [14] k. simon, february 2022.”digital 2022: kenya” [online]. available: https://datareportal.com/reports/digital-2022-kenya. [15] m. jakonen et al., "towards an economy of encounters? a critical study of affectual assemblages in coworking,," scandinavian journal of management,, vol. 33, no. 4 https://doi.org/10.1016/j.scaman.2017.10.003., pp. 235-242, 2017. [16] c. spinuzzi et al, “coworking is about community”: but what is “community” in coworking?," journal of business and technical communication,. [17] j. andrea and z. yingqin , "a spatial persspective of innovation and development: innovation hubs in zambia and the uk," royal holloway, university of london, 2011. [18] d. peter et al.,"silicon savannah: the kenya ict services cluster," microeconomics of competitivenessspring 2016, april 2016 [19] o.el suheil, “silicon savannah: tapping the potential of africa’s tech hub”, september 2021. [online]. [20] w. k lilian. and k.n isaac, "teaching note: case 9: m-pesa: a renowned disruptive innovation from kenya," instructors manual for strategic marketing case in emerging markets, pp. 75-77, may 2017. [21] ubuntu, “welcome to silicon savannah: how kenya is becoming the next global tech hub”, may 2022. [online]. available: https:www.ubuntu.life/en-ke/blogs/news/welcome-to-the-silicon-savannah-how-kenya-is-becoming-thenext -global-tech-hub. [accessed 21 november 2022]. [22] afrilabs, "who we are," immrsv africa, 2021. [23] j. omena, may 2022. “afrilabds and association of countrywide innovation hubs, kenya sign mou to support grassroot ubs [online]. https://data.worldbank.org/indicator.it.net.user.zs?end=2020&locations=ke&start=1990&view=chart https://www.itu.int/hub/2021/11/bridging-the-digital-divide-with-innovative-finance-and-business-models/ vol. 4, no.2 july 2023 | 62 [24] j amina, "kenya's konza techno city: utopian vision meets social reality," independent study project (isp) collection.2024, 2015. [25] g.o kenya, 2022, discover konza technopolis-a global tecnology and innovation hub.”, 2022. [online]. available: https://konza.go.ke. [accessed 3 december 2022]. [26] c. mwoki, "over 300 innovation hubs operationalized," kenya news agency, 2019. [online]. [27] n moses, "ochieng: this is how ajira digital initiative has benefited the youth," nation media group, 2022. [online]. [28] j. tom, “kenya’s bungoma county to launch $2m tech hub”, november 2015. [online]. available: http://disrupt-africa.com/2015/11/04/kenyas-bungoma-county-to-launch-2m-tech-hub/. [29] rvic, “powering ideas into solutions”, 2022. [online]. available: http://rvic.co.ke. [accessed 11 december 2022]. [30] un, “information and communication technologies (icts)”, 2022. [online]. https://konza.go.ke/ http://rvic.co.ke/ vol. 5, no.2 june 2024 | 85 early breast cancer detection in coimbra dataset using supervised machine learning (xgboost) ahmed sami jaddoa business informatics college, university of information technology and communications, iraq email: ahmed.sami@uoitc.edu.iq abstract worldwide, breast cancer (bc) represents one of the serious health concerns for adult females. the early detection and accurate prediction of risks are vital for the provision of optimum care and enhancement of patient outcomes. in the past few years, promising large data merging and ensemble learning algorithms appeared for the purpose of classification and prediction of bc risk. in the area of medical applications, methods of machine learning (ml) are crucial. early diagnosis is necessary for a more efficient carcinoma treatment. this study’s aim is to classify the carcinoma with the use of the 10 predictors that are found in breast cancer coimbra dataset (bccd). presently, early diagnoses are necessary. the rates of cancer survival could be raised in the case where it is discovered early. methods of machine learning offer effective way for data classifying and making early disease diagnoses. this study utilizes bccd for the classification of bc cases utilizing xgboost algorithm. based on performance criteria, early detection of bc is the primary goal. the xgboost classifier in this research achieved 98% precision, 98.32% accuracy, 99% f1-score, and 97% recall. keyword: machine learning, xgboost, z-score, bccd. 1. introduction humans are more vulnerable than ever to many forms of cancer in the last few years. an estimated one in six deaths globally result due to cancer, which makes it one of the top death causes globally. bc is the most prevalent type of cancer in terms of newly diagnosed cases. about 40,920 women died from bc alone in the year 2018. the world health organization (who) estimated that 2.90 million women receive a bc diagnosis annually. no less than 100 diseases that affect various bodily parts are referred to be cancer [1]. the most prevalent cancer all over the world is bc. in the year 2020, there will likely be no less than 2.26 million new cases of bc, according to who research on the disease's current and prospective impact [2]. the retrieval and maintenance of patients' electronic medical records and related devices are examples of healthcare technology. it has never been easy to diagnose and treat hematological diseases when cancer is present. nowadays, a staggering portion of the populace suffers from one or more diseases. medical research has made enormous strides in the last few years. even with such advancements, the general population still knows incredibly little about disease and health. it's possible that a sizable section of the populace has health problems, some of which could be lethal [3]. through creating a prediction model, it is possible to diagnose diseases early on and provide patients with more effective care. in earlier research, ml-based models were employed to identify bc, and they show noteworthy efficacy [4]. one aspect of ai that enables the system to obtain information without explicit expertise is ml. supervised algorithms are employed in many classification applications because they leverage human outputs and inputs to improve prediction accuracy and streamline the training process. as a result, the use of ml in healthcare has expanded [5][6]. ml is p-issn: 2715-2448 | e-issn: 2715-7199 vol.5 no.2 june 2024 buana information technology and computer sciences (bit and cs) mailto:ahmed.sami@uoitc.edu.iq vol. 5, no.2 june 2024 | 86 emerging as a major diagnostic tool for patients in the medical field. in the case when a task is large and difficult to program, ml is used as an analytical technique. examples of such tasks include anticipating pandemics, evaluating genomic data, and turning medical records into knowledge [7]. 2. literature review for classifying bc patients, many different kinds of studies were done recently. research by sakri et al. [8] looked into a number of data mining (dm) methods to predict the recurrence of bc. they employed particle swarm optimization (pso) as feature selection technique for k-nearest neighbor (k-nn), naïve bayes (nb), and rapid decision tree (dt) learners in order to increase the accuracy of the prediction model. the results of this study have shown that the prediction model of the recurrence of bc through using fast decision tree learner (reptree) as classifier had higher accuracy rate, which has been equal to 76.3i% in comparison to k-nn and nb classifiers. keles etal. [9] used non-invasive and painless approaches for bc prediction and diagnosing with the use of dm algorithms. they have studied several algorithms for bc classification utilizing ten-fold cross-validation method for the assessment of each one's predictive power for the relative results, using readings from antenna as a dataset. the optimal algorithms were ibk, bagging, random committee, rf, and simple classification and regression tree (simplecart), which had an over 90% detection accuracy. through using patient data that is related to bc, ferroni etal. [10] have made an identification of ml-based decision support system (dss) significance combined with the random optimization (ro). the dss model was built with the use of multiple kernel learning (mkl), which has first been created for assessing risks of cancer-related thrombosis. the mkl was expanded for the prediction of disease progression risks for patients with breast cancer in oncology setting. as it has been foreseen, their suggested model had produced a 10.90 hazard ratio (hr) and a c-index for pfs (i.e., progression-free survival) of 0.84, with a 86% rate accuracy. akben [11] had presented a decision tree model for the diagnosis of bc by utilizing coimbra data-set. this dt had employed gini index in their research for ascertaining attribute importance degree. compared with existing models, which include k-nn, ann, svm, nb, adaptive boosting (adaboost), and so on, the proposed diagnostic approach had a 90.52% accuracy rate, according to results. an ensemble model for the prediction of breast cancer has been applied by nanglia et al. [12] to the coimbra data-set. in this study, stacking has been utilized for the construction of ensemble model that consists of the three ml algorithms svms, dt, and knn. comparing their suggested ensemble model to other classifiers they employed in the research; it achieved the highest accuracy score of up to 78%. moreover, the chi-square approach was used for the determination of the top five features, which included glucose, insulin, bmi, resistin, and homeostasis model assessment (homa) values. vol. 5, no.2 june 2024 | 87 3. materials and methods fig.1. process of bc detection a. dataset and attributes in the case when related works have been analyzed, various approaches were utilized to diagnose bc. multiple datasets are accessible for the identification of bc. uci ml repository provided the bccd, which consists of 4000 observations and 10 attributes, one of which is a class variable (1 = healthy, 2 = patient). the dataset's attributes are listed in table 1. table1. data-set attributes no attribute name type 1 2 3 4 5 6 7 8 9 10 age bmi glucose insulin homa leptin adiponectin resistin mcp.1 classification numeric numeric numeric numeric numeric numeric numeric numeric numeric numeric b. dataset preprocessing one of the most important phases of ml classification is data preprocessing, since cleaner data typically yields higher classification results. the model's pre-processing methods are outlined as follows: reduce noisy data: in ml, attribute noise as well as class noise are the two types of data noise. nonetheless, attribute noise is decreased in the suggested model to improve accuracy. the data-set is split up into sections for process analysis training and testing. thirty percent is utilized for testing and seventy percent is used for training. c. classification through classifying the records into predefined classes, classification is a model utilized for predicting the future behavior regarding the data. with the use of testing and training data sets, accurate disease detection is possible in classification. to construct the prediction, it suggests using two ml models. test data is used after training data across ml models, i.e., xgboost predicts using each learned model individually. among the criteria are recall, f1-score, accuracy, and precision. vol. 5, no.2 june 2024 | 88 a. extreme gradient boosting (xgboost) classifier xgboost can be defined as a gradient-boosting method that makes use of dts and ensemble ml methods. this ml method can yield very important results because of its scalability and high processing speed. xgboost is applied to both classification and regression tasks. the concept of this method is to find weak classifiers iteratively in order to get a correct classification. it creates customized dts through the gradient descent technique through first establishing a range of threshold values, which are after that updated repeatedly through reducing residuals throughout tree construction. regression trees can be considered as the weak learners when gradient boosts are applied, and each one maps an input dataset that shows one of its leaves contains a continuous mark. regularized function (l1 & l2) that is minimized is a convex loss function that is based upon the difference between the goal output and predictions. in order to predict the mistakes or residues of earlier trees, the training iteratively adds new trees that are subsequently integrated with the earlier trees to provide the final prediction [13]. b. performance metric evaluation method model performance is assessed using the accuracy, f1-score, recall, and precision metrics. eqs. (1) and (2) characterize the accuracy value regarding the classification model, respectively; comparably, eqs. (3) and (4) denote the mathematical formula of recall and the f1-score, respectively [14]. accuracy = 𝑇𝑃+𝑇𝑁 𝑇𝑃+𝑇𝑁+𝐹𝑃+𝐹𝑁 (1) precision = 𝑇𝑃 𝑇𝑃+𝐹𝑃 (2) recall = 𝑇𝑃 𝑇𝑃+𝐹𝑁 (3) f1-score = 2∗(𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛∗𝑅𝑒𝑐𝑎𝑙𝑙) (𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛+𝑅𝑒𝑐𝑎𝑙𝑙) (4) 4. results following the algorithms described in the preceding section, the following outcomes were attained: table 2. performance measure of xgboost classifier classifier accuracy precision recall f1-score xgboost 98.32% 98% 97% 99% 5. conclusion a predictive model for the likelihood of bc development was reported in this paper. the suggested model's study and evaluation demonstrate how effective and simple it is to apply. for patients with classified bc, the xgboost algorithm is used on the bccd. with the use of z-score outlier detection method, one can identify outliers in a clean dataset during data preparation and apply robust scaling to address the outliers. afterward, xgboost algorithm was used. based on the examination of test outcomes, the xgboost algorithm has attained the maximum accuracy of 98.32%. references [1] s. aamir et al., “predicting breast cancer leveraging supervised machine learning techniques,” comput. math. methods med., vol. 2022, 2022, doi: 10.1155/2022/5869529. [2] r. gonzales martinez and d. m. van dongen, “deep learning algorithms for the early detection of breast cancer: a comparative study with traditional machine learning,” informatics med. unlocked, vol. 41, no. april, p. 101317, 2023, doi: 10.1016/j.imu.2023.101317. [3] y. amethiya, p. pipariya, s. patel, and m. shah, “comparative analysis of breast cancer vol. 5, no.2 june 2024 | 89 detection using machine learning and biosensors,” intell. med., vol. 2, no. 2, pp. 69–81, 2022, doi: 10.1016/j.imed.2021.08.004. [4] g. alfian et al., “predicting breast cancer from risk factors using svm and extra-treesbased feature selection method,” computers, vol. 11, no. 9, 2022, doi: 10.3390/computers11090136. [5] a. sami jaddoa, s. j. saba, and e. a.abd al-kareem, “liver disease prediction model based on oversampling dataset with rfe feature selection using ann and adaboost algorithms,” buana information technology and computer sciences (bit and cs), vol. 4, no. 2. pp. 85–93, 2023, doi: 10.36805/bit-cs.v4i2.5565. [6] c. li, y. weng, y. zhang, and b. wang, “a systematic review of application progress on machine learning-based natural language processing in breast cancer over the past 5 years,” diagnostics, vol. 13, no. 3, 2023, doi: 10.3390/diagnostics13030537. [7] ahmed sami jaddoa, “heart disease prediction system using (smote technique),” vol. 050006, 2023. [8] s. b. sakri, n. b. abdul rashid, and z. muhammad zain, “particle swarm optimization feature selection for breast cancer recurrence prediction,” ieee access, vol. 6, pp. 29637–29647, 2018, doi: 10.1109/access.2018.2843443. [9] m. kaya keleş, “breast cancer prediction and detection using data mining classification algorithms: a comparative study,” teh. vjesn., vol. 26, no. 1, pp. 149–155, 2019, doi: 10.17559/tv-20180417102943. [10] p. ferroni, f. m. zanzotto, s. riondino, n. scarpato, f. guadagni, and m. roselli, “breast cancer prognosis using a machine learning approach,” cancers (basel)., vol. 11, no. 3, pp. 1–9, 2019, doi: 10.3390/cancers11030328. [11] s. b. akben, “determination of the blood, hormone and obesity value ranges that indicate the breast cancer, using data mining based expert system,” irbm, vol. 40, no. 6, pp. 355–360, 2019, doi: 10.1016/j.irbm.2019.05.007. [12] s. nanglia, m. ahmad, f. ali khan, and n. z. jhanjhi, “an enhanced predictive heterogeneous ensemble model for breast cancer prediction,” biomed. signal process. control, vol. 72, no. july 2021, p. 103279, 2022, doi: 10.1016/j.bspc.2021.103279. [13] m. m. hassan et al., “a comparative assessment of machine learning algorithms with the least absolute shrinkage and selection operator for breast cancer detection and prediction,” decis. anal. j., vol. 7, no. may, p. 100245, 2023, doi: 10.1016/j.dajour.2023.100245. [14] z. salod and y. singh, “comparison of the performance of machine learning algorithms in breast cancer screening and detection: a protocol,” j. public health res., vol. 8, no. 3, pp. 112–118, 2019, doi: 10.4081/jphr.2019.1677. vol. 5, no.2 june 2024 | 99 quantum computing’s paradigm shift: implications and opportunities for cloud computing veera brahmam ainala1, yaswanth arikatla2, varma seru3, sahil dasu4, b. bikram5 department of computer science and engineering, koneru lakshmaiah education foundation,vaddeswaram, andhra pradesh, india12345 email: veerabrahmam401@gmail.com1, yaswanth.arikatla56@gmail.com2, varmaseru2123@gmail.com3, iamsahil441@gmail.com4, basaba.bikram@kluniversity.in5 abstract the field of computing is about to undergo a revolution thanks to quantum computing, a ground-breaking innovation based on the ideas of quantum physics. the enormous implications and potential that quantum computing has for cloud computing are explored in this publication. we examine the difficulties faced by current cryptography systems as well as prospective improvements in fields like machine learning, simulations, and data security as we delve into the basic alterations in computing paradigms. additionally, we go over the revolutionary possibilities for cloud computing, such as the creation of quantum-safe cloud security solutions, hybrid computing models, and quantum cloud services. keyword: quantum-computing, cloud computing, quantum cloud services, hybrid cloud services, quantum cryptography. i. introduction in the realm of computing, a revolutionary shift is on the horizon, one that promises to alter the very fabric of how we process information, solve problems, and secure data. quantum computing, founded upon the principles of quantum mechanics, stands as a testament to human ingenuity, challenging the traditional boundaries of classical computing [1]. quantum computing harnesses the unique behavior of quantum bits, or qubits, which can exist in multiple states simultaneously due to the phenomena of superposition and entanglement [2]. this ability to process vast amounts of information in parallel presents a paradigm shift that holds the potential to transform various industries, including the landscape of cloud computing. traditionally, classical computing, relying on binary bits representing 0s and 1s, has been the backbone of our digital age. however, as the complexities of problems in fields such as cryptography, optimization, and simulations continue to grow, classical computers face insurmountable challenges in terms of processing power and time efficiency [3]. quantum computing emerges as the beacon of hope, offering exponential computational advantages for specific problem sets. the capabilities of quantum computers, when harnessed effectively, are poised to reshape the future of cloud computing, unlocking new horizons of possibility and efficiency. this paper explores the profound implications and the vast array of opportunities that the fusion of quantum computing and cloud computing presents. it delves into the vulnerabilities of classical cryptographic systems, the transformative potential of quantum simulations, and the innovative applications of quantum-enhanced machine learning algorithms within cloud environments [4]. furthermore, it examines the advent of quantum cloud services and the integration of quantum-safe security protocols into cloud architectures [5]. through a comprehensive analysis of the synergies between quantum and cloud computing, this paper sheds light on the intricate interplay between these two transformative technologies. p-issn: 2715-2448 | e-issn: 2715-7199 vol.5 no.2 june 2024 buana information technology and computer sciences (bit and cs) mailto:veerabrahmam401@gmail.com1 mailto:yaswanth.arikatla56@gmail.com2 mailto:varmaseru2123@gmail.com3 vol. 5, no.2 june 2024 | 100 fig 1. structure of quantum computing fig. 2. representation of difference between classical and quantum computer quantum computing, a frontier technology, harnesses the peculiar principles of quantum mechanics to perform computations at a scale previously thought impossible. at its core, quantum computing operates on qubits, the quantum counterpart to classical bits. unlike classical bits, which are either 0 or 1, qubits can exist in multiple states simultaneously due to a phenomenon called superposition[1]. this property allows quantum computers to process vast amounts of information in parallel, offering unprecedented computational power. a. superposition and qubits: in classical computing, a bit can either be 0 or 1, representing two distinct states. however, qubits exist in a superposition of states, meaning they can be both 0 and 1 at the same time. this inherent duality exponentially increases the computational possibilities of quantum systems. when qubits are entangled, the state of one qubit instantaneously influences the state of another, regardless of the distance between them, a phenomenon essential for quantum computing [2]. vol. 5, no.2 june 2024 | 101 fig 3. superposition and qubits representation b. quantum gates and quantum circuits: quantum computations are executed using quantum gates, which manipulate qubits' states. these gates perform operations such as flipping the qubit's state, creating entanglement, or applying complex mathematical transformations. quantum gates are combined to form quantum circuits, which represent the sequence of operations applied to qubits during a computation. the specific arrangement of gates in a quantum circuit determines the output, and due to quantum parallelism, the quantum computer evaluates all possible paths simultaneously [3]. fig 4. quantum gates and circuits c. quantum algorithms: quantum algorithms leverage the unique properties of qubits to solve problems exponentially faster than classical algorithms. for instance, shor's algorithm, a groundbreaking quantum algorithm, factors large integers exponentially faster than the best-known classical algorithms. another significant algorithm is grover's search algorithm, which performs unstructured searches quadratically faster than classical counterparts [4]. vol. 5, no.2 june 2024 | 102 fig 5. quantum algorithm d. quantum computing and quantum cryptography: quantum computing also intersects with quantum cryptography, a field focused on secure communication. quantum key distribution (qkd) protocols like bb84 use quantum properties to ensure the secrecy of encryption keys. qkd relies on the principle of quantum indeterminacy: any attempt to eavesdrop on the quantum channel irreversibly alters the quantum state, alerting the communicating parties to potential tampering [5]. fig 6. quantum cryptography structure and process a. implications for cloud computing a. cryptographic vulnerabilities quantum computing poses a significant threat to classical cryptographic systems, which rely on the difficulty of certain mathematical problems for their security. one of the most prominent algorithms in this context is shor's algorithm, devised by mathematician peter shor. shor's algorithm efficiently factors large integers, breaking rsa encryption and related protocols, which are widely used for securing data transmission [6]. the implication of this breakthrough is profound; it renders data encrypted with current public-key cryptography vulnerable to decryption once large-scale, practical quantum computers become available. b. quantum simulations quantum simulations are another domain where quantum computing presents transformative opportunities, particularly for scientific research. quantum systems are inherently complex and difficult to simulate using classical computers, especially when dealing with large-scale quantum phenomena. quantum computers, on the other hand, can accurately model the behavior of quantum systems, providing insights into areas such as material science, drug discovery, and fundamental physics [7]. vol. 5, no.2 june 2024 | 103 these simulations can be integrated into cloud computing environments to facilitate real-time data analysis and experimentation. cloud platforms equipped with quantum simulation capabilities enable researchers to run complex simulations without the need for massive local computational resources[8]. this integration empowers scientists to explore intricate quantum phenomena, accelerating the pace of discovery and innovation. c. integration into cloud computing real-time data analysis: quantum simulations integrated into cloud platforms enable real-time analysis of complex data streams. this capability is invaluable for fields like financial modeling, where rapid analysis of market data and risk assessment are critical[9]. drug discovery: quantum simulations assist in simulating molecular interactions accurately. cloud platforms with quantum simulation capabilities accelerate drug discovery by predicting molecular behaviors and interactions, aiding in the development of new pharmaceuticals[10]. incorporating these quantum simulations into cloud computing not only enhances the efficiency of computations but also democratizes access to advanced scientific tools, fostering innovation and exploration in various fields. fig 7. cloud integration among different services. fig 8. cryptographic vulnerabilities b. opportunities for cloud computing a. quantum cloud services cloud service providers have recognized the potential of quantum computing and are investing in quantum cloud services, allowing businesses to access quantum power without the substantial costs of building and maintaining quantum hardware. these services offer cloud-based access to quantum computers, enabling companies to experiment, develop, and run quantum algorithms without the need for in-house quantum expertise. ibm quantum experience and amazon braket are examples of cloudbased platforms providing access to quantum computing resources, fostering innovation without the prohibitive expenses [11]. vol. 5, no.2 june 2024 | 104 b. quantum-safe cloud security as quantum computers threaten classical encryption methods, quantum-safe cloud security has become a paramount concern. researchers are developing quantum-resistant encryption algorithms, also known as post-quantum cryptography, which are believed to be secure against attacks from quantum computers. these algorithms are being integrated into cloud security protocols to ensure the confidentiality and integrity of data in the quantum era. nist (national institute of standards and technology) is actively involved in standardizing post-quantum cryptographic algorithms to prepare for the advent of quantum computers [12]. c. quantum-inspired algorithms quantum-inspired algorithms are classical algorithms that draw inspiration from quantum computing principles, such as superposition and entanglement, to solve problems more efficiently. these algorithms are designed to mimic quantum behavior and have shown promising results in various fields. for instance, quantum-inspired algorithms have been applied in optimization problems, machine learning, and complex data analysis. integrating these algorithms into cloud computing environments enhances classical computing capabilities, enabling faster and more accurate solutions to complex problems [13]. fig 9. quantum algorithms for different problems. c. case studies and practical applications a. real-world examples one notable example of a business harnessing quantum cloud services is daimler ag, the automotive giant. daimler collaborated with ibm to explore the potential of quantum computing in optimizing traffic flow and reducing congestion. by utilizing ibm's quantum cloud services, daimler's researchers were able to model intricate traffic patterns more accurately. this collaboration allowed daimler to analyze and process vast amounts of real-time traffic data efficiently, leading to the development of innovative traffic management strategies [14]. fig 10. real world leap in quantum computing vol. 5, no.2 june 2024 | 105 b. success stories one compelling success story in the realm of quantum and cloud computing integration is the research conducted by the google quantum ai lab. google's team utilized hybrid cloud models, combining classical and quantum computing resources, to simulate the behavior of quantum systems. by leveraging cloud-based quantum simulators alongside classical computing infrastructure, they overcame challenges related to quantum noise and decoherence. this integration led to breakthroughs in understanding quantum systems, enhancing google's capabilities in quantum algorithm research [15] fig 11. process for demonstrating quantum supremacy. fig 12. number of qubits, n vol. 5, no.2 june 2024 | 106 fig 13. heat map showing single(e1; crosses) and two-qubit (e2; bars) pauli errors for all qubits operating simultaneously. the layout shown follows the distribution of the qubits on the processor. ii. literature work the process of our research in quantum and cloud computing domain is shown in fig. 1. we have collected 50 research articles from different well-known sites such as ieee xplore, science direct, tech science for our topic. these literature works form the foundation for understanding the synergy between quantum computing and cloud computing, providing insights into theoretical principles, practical applications, and the societal impact of these technologies. table 1. a study on existing theories on quantum and cloud computing s. no author & year methodology remarks 1 nielsen & chuang (2010) theoretical exploration of quantum computing principles and algorithms. fundamental textbook providing detailed insights into quantum computation and information. 2 rieffel & polak (2011) conceptual explanation of quantum computing, focusing on approachable language and illustrations. beginner-friendly introduction to quantum computing concepts and their applications. 3 mcmahon (2008) explains quantum computing concepts and algorithms using accessible language and examples. provides a clear understanding of quantum computing, making it accessible to readers with varied backgrounds. 4 yanofsky & mannucci (2008) theoretical exploration of quantum algorithms and complexity theory. geared towards computer scientists, delves into the theoretical aspects of quantum computing and its algorithms. 5 erl, mahmood & puttini (2013) covers cloud computing concepts, technologies, and architectural principles. comprehensive guide to understanding cloud computing, laying the foundation for cloud-based quantum computing discussions. 6 preskill (2018) theoretical exploration of noisy intermediate-scale quantum (nisq) devices and their potential applications. discusses the current state of quantum computing, focusing on practical implementations and potential applications. 7 zhong et al. (2020) experimental demonstration of quantum computational advantage using photons. presents experimental results showcasing the practical progress made in achieving quantum computational advantage. 8 berta et al. (2020) exploratory research on the societal and policy implications of quantum technologies. discusses the broader societal impact of quantum technologies, highlighting policy considerations and security implications. vol. 5, no.2 june 2024 | 107 after comprehending the abstract, we reduced the articles from 100 to 45, then after studying various quantum implications and algorithms, we reduced them to 8, as shown in table 1. following the literature work, we understand the different cloud security, algorithms, and quantum consequences. iii. conclusion in this journal, we explored the symbiotic relationship between quantum computing and cloud computing, highlighting the transformative potential of their integration. quantum computing, with its unique principles of superposition and entanglement, enhances computational capabilities, while cloud computing provides the necessary infrastructure and accessibility. we discussed how quantum cloud services, hybrid models, quantum-safe cloud security, and quantum-inspired algorithms are reshaping various industries, from traffic management to scientific research. future developments the future of quantum computing holds exciting possibilities. anticipated developments include the scaling up of quantum hardware, improving qubit coherence times, and advancing error correction techniques. these advancements will enable the development of more stable and powerful quantum computers. additionally, the field of quantum machine learning is poised to grow, integrating quantum algorithms with classical machine learning techniques for unprecedented data analysis capabilities. quantum communication technologies, such as quantum key distribution, are also expected to mature, revolutionizing secure data transmission in the cloud environment [16]. references [1] nielsen, m. a., & chuang, i. l. (2010). quantum computation and quantum information. cambridge university press. [2] preskill, j. (1998). quantum computation and quantum information. caltech, pasadena, ca, usa, 43(2), 42-48. [3] mermin, n. d. (2007). quantum computer science: an introduction. cambridge university press. [4] shor, p. w. (1994). algorithms for quantum computation: discrete logarithms and factoring. proceedings 35th annual symposium on foundations of computer science, 124-134. [5] bennett, c. h., & brassard, g. (1984). quantum cryptography: public key distribution and coin tossing. proceedings of ieee international conference on computers, systems and signal processing, 175-179. d. zhang et al., "research on diagnosis characteristics of wheat powdery mildew under different severity grading standards," 2019 8th international conference on agrogeoinformatics (agro-geoinformatics), pp. 1-6, 2019. [6] shor, p. w. (1994). algorithms for quantum computation: discrete logarithms and factoring. proceedings 35th annual symposium on foundations of computer science, 124-134. [7] peruzzo, a., mcclean, j., shadbolt, p., yung, m. h., zhou, x. q., love, p. j., ... & o'brien, j. l. (2014). a variational eigenvalue solver on a photonic quantum processor. nature communications, 5, 4213. [8] wecker, d., hastings, m. b., & troyer, m. (2015). progress towards practical quantum variational algorithms. physical review a, 92(2), 022305. [9] rebentrost, p., mohseni, m., & lloyd, s. (2014). quantum support vector machine for big data classification. physical review letters, 113(13), 130503. [10] mcclean, j. r., romero, j., babbush, r., aspuru-guzik, a. (2016). the theory of variational hybrid quantum-classical algorithms. new journal of physics, 18(2), 023023. [11] ibm quantum experience: https://quantum-computing.ibm.com/ [12] amazon braket: https://aws.amazon.com/braket/ [13] national institute of standards and technology (nist). (2021). post-quantum cryptography standardization. https://csrc.nist.gov/projects/post-quantum-cryptography [14] vadhan, s. (2018). the theory of quantum computing. notices of the ams, 65(2), 203-215. vol. 5, no.2 june 2024 | 108 [15] ibm quantum. (2019). ibm q network strengthens research on traffic optimization. https://www.research.ibm.com/ibm-q/network/academic-partnerships/case-studies/daimler.shtml [16] google ai blog. (2018). achieving quantum supremacy. https://ai.googleblog.com/2019/10/quantum-supremacy-using-programmable.html [17] preskill, j. (2018). quantum computing in the nisq era and beyond. quantum, 2, 79. zhong, h. s., et al. (2020). quantum computational advantage using photons. science, 370(6523), 1460-1463. [18] rieffel, e., & polak, w. (2011). quantum computing: a gentle introduction. the mit press. [19] mcmahon, d. (2008). quantum computing explained. john wiley & sons. [20] yanofsky, n. s., & mannucci, m. a. (2008). quantum computing for computer scientists. cambridge university press. [21] erl, t., mahmood, z., & puttini, r. (2013). cloud computing: concepts, technology & architecture. prentice hall. [22] berta, m., et al. (2020). quantum technologies: impact on policy, security, and society. science robotics, 5(41), eaba4564. [23] quantum cloud computing: how does it fit into hpc ecosystem? by h. a. khadem and k. s. trivedi (2016). [24] quantum-assisted cloud computing by r. wille and r. drechsler (2016). [25] secure cloud computing with quantum key distribution by m. al azwari, a. shahrabi, and k. k. r. choo (2018). [26] quantum computing for enhanced machine learning by a. maria, v. babovic, and s. pallickara (2017). [27] cloud quantum computing of an atomic nucleus by g. d. barron et al. (2018). [28] quantum cloud services for science and society by i. l. chuang, s. j. herbert, n. m. linke, et al. (2018). [29] quantum blockchain using entanglement in time by g. karakonstantis, i. koutsopoulos, and k. g. saurabh (2019). [30] quantum-safe key management for cloud-based systems by a. pawlowski, m. a. adnan, and k. k. r. choo (2019). [31] quantum machine learning: a classical approach by s. lloyd, m. mohseni, and p. rebentrost (2013). [32] quantum machine learning in feature hilbert spaces by m. schuld, i. sinayskiy, and f. petruccione (2015). [33] towards quantum-safe cloud computing: a survey by a. rasoolzadegan, a. s. ibrahim, and j. markarian (2018). [34] quantum cloud computing: a survey by d. thapliyal and a. r. calderbank (2017). [35] quantum machine learning: a review by x. xu, j. sun, and j. du (2020). [36] quantum technologies: an old new story by r. ursin et al. (2020). [37] quantum computing and the ultimate limits of computation: the case for a national investment by e. farhi and a. w. harrow (2016). [38] secure quantum cloud computing by y. li and z. xu (2019). [39] quantum cryptographic algorithms for cloud security by x. fu et al. (2017). [40] quantum cloud computing: a new epoch by l. lu et al. (2018). [41] quantum computing for communications: an overview and advances by t. wang and y. cai (2018). [42] quantum-assisted cloud computing: is it feasible by f. yao et al. (2016). [43] quantum computing in the cloud with ibm q experience by a. mezzacapo et al. (2019). [44] quantum computing in the cloud with ibm q experience by a. mezzacapo et al. (2019). [45] quantum-assisted cloud-edge computing for the internet of things by l. wang, q. zhu, and d. k. y. yau (2019). [46] quantum cloud computing: challenges and opportunities by z. shen, et al. (2022) vol. 5, no.2 june 2024 | 109 p-issn : 2715-2448 | e-isssn : 2715-7199 vol.3 no.1 january 2022 buana information tchnology and computer sciences (bit and cs) 11 | vol.3 no.1 january 2022 design of customer satisfaction application at bca kcp rengasdengklok using c.45 algorithm method april lia hananto 1 study program information system faculty of engineering and computer science, universitas buana perjuangan karawang aprilia@ubpkarawang.ac.id shofa sofiah hilabi 2 study program information system faculty of engineering and computer science, universitas buana perjuangan karawang shofa.hilabi@ubpkarawang.ac.id ‹β› detrie noviani 3 study program information system faculty of engineering and computer science, universitas buana perjuangan karawang detrienoviami@mhs.ubpkara wang.ac.id abstrak—bank bca selalu meningkatkan kualitas layanan sesuai dengan slogan “senantiasa disisi anda”. penilaian tersebut terdiri dari 4 atribut (waktu, akurat, fokus dan kepuasan). setiap atribut memiliki bobot nilai angka 1 (sangat tidak puas) sampai 5 (sangat puas). penilian tersebut masih dilakukan secara manual (menggunakan kertas), oleh sebab itu penulis pada penelitian ini menggunakan algoritma c.45 dengan dilakukan 3 kali pengujian sehingga menghasilkan klasifikasi yang diperoleh bahwa nilai akurasi yaitu mencapai 88,75% dengan nilai auc yaitu 0, 744 dan pengujian pada aplikasi yang dibuat menghasilkan 0,722. dapat disimpulkan bahwa penilaian layanan di bca kcp rengasdengklok termasuk kelompok klasifikasi yang cukup baik dikarenakan nilai auc-nya antara 0.70-0.80. kata kunci: pelayanan, algoritma c4.5, klasifikasi abstract—bank bca continuously improves service quality following the slogan “always by your side”. the assessment consists of 4 attributes (time, accuracy, focus, and satisfaction). each feature has a weighted value of 1 (very dissatisfied) to 5 (delighted). the assessment is still done manually (using paper). therefore the authors in this study used the c.45 algorithm with three tests carried out to produce a classification obtained that the accuracy value reached 88.75% with an auc value of 0.744, and testing on the application that was made resulted in 0.722. it can be concluded that the service assessment at bca kcp rengasdengklok belongs to a reasonably good classification group because the auc value is between 0.70-0.80. keywords: service, c4.5 algorithm, classification i. introduction banking, commonly referred to as a bank, is a business entity that provides financial services for all levels of society. according to law number 10 of 1998, "bank is a business entity that collects public funds in the form of savings and distributes them to the public in the form of credit and or other forms to improve people's living standards [1]. bank bca is the largest private bank in indonesia, established in 1957. in providing services to the banking world, it always wants to provide the best service for customers to maintain and for the long-term sustainability of a company. customers judge the service on what they receive with what they expect [2]. the way to provide the best service to customers is to establish good relationships with customers and accept complaints felt by customers to improve service quality. the hope of every bank is customer satisfaction to develop the long-term sustainability of a company. as bank bca's commitment is "always by your side". the benchmark for the service quality of bca kcp rengasdengklok there is four attributes, namely as follows: time (time), focus (focus), accuracy (accurate), and satisfaction (satisfaction). of the four attributes above are mainstays that must be done to improve service quality at bca kcp rengasdengklok. in improving the quality of service to customer satisfaction, when the customer finishes a transaction, several questions and answers are given in numbers 1, namely very dissatisfied (stp), to 5, namely very satisfied (sp). this assessment data must be managed properly by a web-based system to make it easier for customers to provide assessments and make it easier for each employee to make reports and find out the quality of services that have been provided. the procedure that is running at the branch at the time of providing an assessment by customers is still manual, namely using paper by filling out a survey form manually, where the data is in the form of an archive so that it is vulnerable to damage and loss, therefore it is necessary to have an application with better storage (database) and safe. the reports generated on the service assessment are not timely (real-time). therefore it takes a long time to find out the results of the report. meanwhile, to analyze and manage the data using data mining to produce an information [3]. in analyzing this research, the information obtained uses the classification and calculation methods algorithm c4.5 to produce a decision tree to determine the level of service satisfaction. ramadhan, in his research using the c.45 algorithm method with rapidminer tools, produces an accuracy value of 96.50% [4]. 12 | vol.3 no.1, january 2022 meanwhile, febriyanto [5] and dhika [6] their research using data mining classification resulted in the same accuracy value, namely 91%. in addition, [7] also tested customer satisfaction and obtained 93.33% accuracy results with the excellent classification criteria in the confusion matrix. not only determining customer satisfaction but the c4.5 algorithm method can also be used in determining the decisions of scholarship recipient students, resulting in a decision tree, namely 14 students who deserve scholarships [8]. as for determining the satisfaction of brt bus passengers using the c4.5 algorithm method, the accuracy results are 95%, indicating that the satisfaction of the brt bus is very good [9]. thus, the study shows that the c4.5 algorithm is very suitable for measuring customer/customer satisfaction. this study examines customer satisfaction regarding the services that have been provided. to determine the level of customer satisfaction using the c4.5 algorithm method, the c4.5 algorithm calculation using rapidminer tools and programming languages in building applications, namely php and mysql as the database. based on the explanation of the problem, the authors chose the title "designing customer satisfaction applications at bca kcp rengasdengklok using the c4.5 algorithm method" [5]. ii. method a. bank the definition of a bank as regulated in law number 10 of 1998 concerning banking is "a business entity that collects public funds in the form of savings and distributes them to the public in the form of credit and or other forms in order to improve the standard of living of the community.[1] the definition of a bank according to another opinion is an institution in carrying out financial activities needed by the community. [10] b. service according to the big indonesian dictionary, it is stated: "services are matters and conveniences provided in connection with" buying and selling goods and services” [11] according to tjiptono in the journal [12] “consumer satisfaction as a conscious evaluation or cognitive assessment concerns whether the product performance is relatively good or bad or whether the product is suitable or not suitable for the intended use”. meanwhile, according to another opinion, service is an attitude that is expected and results in the reality of a consumer's desire so as to form a service quality. [2]. c. customer satisfaction consumer satisfaction is very important because it can be profitable and makes competition within a company [7]. meanwhile, according to kotler, consumer satisfaction is a feeling of the expected reality [2]. d. data mining data mining is a term used to describe the discovery of knowledge in databases" [13]. meanwhile, according to pramudiono, namely "automatic analysis of large or complex data with the aim of finding important patterns or trends that are usually not aware of their existence" [14]. another opinion reveals that data mining is a technique in learning to analyze and knowledge [15]. e. classification of data mining classification is a process of grouping data that will be used to predict data in a decision tree that does not yet have a certain data class [5]. the steps in preparing kdd, namely: data cleaning, data integration, data selection, data transformation, data mining, pattern evaluation, knowledge presentation. the levels of classification in data mining (gorunescu, 2011) are: a. 0.90-1.00 = very good b. 0.80-0.90 = good c. 0.70-0.80 = enough d. 0.60-0.70 = bad e. 0.50-0.60 = wrong f. algoritma c4.5 decisions so as to produce the best and significant accuracy values. meanwhile, according to [3] the c4.5 algorithm is the development of ide3 to create a decision tree with processed data. the first step is to calculate the entropy, the formula is as follows: 𝑛 𝐸𝑛𝑡𝑟𝑜𝑝𝑦(𝑆) = ∑ − 𝑝𝑖 𝑙𝑜𝑔2 𝑝𝑖 𝑖=1 description: s = case set n = number of partitions s pi = the proportion of si to s to make a decision tree, you must first choose the root by determining the highest gain value, the formula is: description: s = case set a = feature n = number of attribute partitions a |si| = proporsi the proportion of si to s |s| = number of cases in s in this study there is a flow of research methods in data collection and system development. the flow of the research method specified in the waterfall method is as follows: 13 | vol.3 no.1, january 2022 figure 2 research method flow iii. results and discussion data processing in designing customer satisfaction applications to get accurate results is carried out with 2 tests, namely manual testing using rapidminer tools and testing using a system designed using the php programming language with a codeigniter framework and a database using mysql. in the depiction of the design using uml modeling which consists of use case diagrams, activity diagrams, class diagrams and sequence diagrams. the flowmap of the proposed system for customer satisfaction applications is: figure 3 flowmap system 1. data collection in the process of data collection will be explained to design and develop the system in this study. data collection is done by using stages of literature study, observation, interviews and questionnaires 2. system design the following is a design model for customer satisfaction applications: 3. use case diagram the function of the use case diagram is to describe the usability of a system. the following is a proposed use case diagram for a service quality application: figure 4 usecase diagram 4. activity diagram the activity is carried out by the user, namely the customer, starting from the user selecting the questionnaire then the system will display the questionnaire, the user can fill in the questionnaire correctly and the system will save the answers from the completed questionnaire. activity diagram for filling the questionnaire as follows: figure 5 activity diagram 5. class diagram class system diagram of the proposed application of customer satisfaction with services at bca kcp rengasdengklok, as follows: 14 | vol.3 no.1, january 2022 figure 6 class diagram 6. sequence diagram sequence diagram describes a process performed by each user as messages sent and received. sequence diagram of the proposed service quality application system, namely: figure 7 sequence diagram 7. system implementation in addition to system design, the interface design for the customer satisfaction application aims to describe an application display design that will be implemented using pencil tools, as follows: figure 8 customer satisfaction survey design implementation of the system in the first stage using rapidminer tools, carried out 3 times of testing. data obtained by 100 respondents by distributing questionnaires on march 17 april 6 2021 through a system that has been created. taken 80 data (calculation of the sample using data solve) as a sample in this research. the data to be processed using rapid miner is as follows: table 1 data that will be imported to rapidminer nama waktu akurat fokus kepuasan layanan polynominal integer integer integer integer binominal id attribute attribute attribute attribute label iqbal 5 5 5 5 satisfied andri mulyadi 5 5 5 5 satisfied vustikawati aulia 5 5 5 5 satisfied abdulmalik 5 5 5 5 satisfied .. .. .. .. .. .. vinka syafana 3 2 2 2 not satisfied after the decision tree processing using rapidminer tools is complete, it will produce a decision tree as follows: figure 9 decision tree results tests carried out 3 timeswith k-fold validation 10, k-fold validation 5 and k-fold validation 3. the results of 3 tests using rapidminer tools are as follows: table 2 results of data analysis using rapidminer k-fold validasi accuracy precision recall auc 10 85,00% 87,88% 93,55% 0,526 5 86,25% 88,06% 95,16% 0,568 3 88,75% 89,55% 96,77% 0,744 based on the table above, the smaller the validation value, the higher the accuracy value obtained. the auc value obtained at validation 3 is 0.744, based on the classification level 0.70-0.80 = enough. it can be concluded that the services at bca kcp rengasdengklok on 17 march – 6 april 2021 are quite satisfactory. based on the results of analysis and testing that produces a decision tree and rules that are formed, the next step is to implement it into a program that has been made using the php programming language with the php framework and mysql database; it looks like this:: 1. login page on the login page for admins and employees, users who will access the application must first log in by entering 15 | vol.3 no.1, january 2022 registered username and password. the display is as follows: figure 10 login page 2. register page for users who do not have an account to login and access the application, they must first register on the register form. display registers are as follows: figure 11 register page 3. customer survey results page customers who have filled out the questionnaire will be stored in the database and displayed in the administrator on the customer survey results menu. the customer survey result page is as follows: figure 12 customer survey results page 4. algoritma c45 page on the c4.5 algorithm page displays the results of entropy, gain and also the conclusion of the assessment results that have been given by the customer. the c45 algorithm page is as follows: figure 13 algoritma c4.5 page figure 14 c4.5 algoritma algorithm results page 5. questionnaire page for customers this questionnaire page is filled out by customers who have transacted at tellers and csos. here's how it looks: figure 15 customer questionnaire page iv. conclusion some conclusions from the research this: 1. the developed web-based application can simplify evaluating staff performance at the contact center of pt. xyz. the calculation method in the application follows the provisions or standards that exist in the internal organization, both the targets to be achieved and the assessment reference. 2. the application developed can help staff monitor work achievement every month. so that if there are poor work achievements in the previous month, staff can find out more quickly and make performance improvements in the following month. in addition, the application can also be helpful for department heads if a special evaluation is needed for staff based on final grades. 16 | vol.3 no.1, january 2022 reference [1] m. abdullah, manajemen dan evaluasi kinerja karyawan. sleman: aswaja pressindo, 2014. [2] d. a. wardani, “pengaruh penerapan aplikasi sistem informasi akuntansi terhadap kinerja karyawan pada pd. bpr rokan hulu pasir penga iran,” j. chem. inf. model., vol. 53, no. 9, pp. 1689–1699, 2017. [3] y. suherman and d. yadewani, “aplikasi sistem informasi penilaian kinerja karyawan,” j-click, vol. 6, no. 2, pp. 201–207, 2019. [4] i. m. hadi, t. tukino, and a. fauzi, “sistem informasi monitoring evaluasi standar pembelajaran menggunakan framework codeigniter,” ciastech 2020, no. ciastech, pp. 443–452, 2020. [5] v. felita, k. saputra, s. keputusan, k. pendidikan, and b. lampung, “aplikasi monitoring kerja karyawan ( e-kinerja ) berbasis web menggunakan framework codeigniter di citra angkasa tercipta ( cat ) bandar lampung,” j. vania, pp. 1–8, 2020. [6] a. marbawi, manajemen sumber daya manusia. lhokseumawe: unimal press, 2016. [7] z. makmur, “pengembangan sistem informasi permintaan pembelian kebutuhan kantor pada dealer management system,” j. teknosain, vol. xv, no. 3, pp. 78–88, 2018. [8] k. yuliana, saryani;, and n. azizah, “percanangan rekapitulasi pengiriman barang berbasis web,” j. sisfotek glob., vol. 9, no. 1, 2019, doi: http://dx.doi.org/10.38101/sisfotek.v9i1.223. [9] r. sabaruddin and w. e. jayanti, jago ngoding pemrograman web dengan php, no. january. surabaya: cv. kanaka media, 2019. [10] i. daqiqil, framework codeigniter sebuah panduan dan best practice. pekanbaru, 2011. [11] b. huda and s. aripiyanto, “aplikasi sistem informasi lowongan pekerjaan berbasis android dan web monitoring (penelitian dilakukan di kab. karawang) 1baenil,” j. buana ilmu, vol. 4, no. 1, pp. 11–24,2019,. [12] r. mall, fundamentals of software engineering fourth edition, 4th ed. delhi: phi learning private limited, 2014. [13] b. huda and b. priyatna, “penggunaan aplikasi content management system (cms) untuk pengembangan bisnis berbasis e-commerce,” systematics, vol. 1, no. 2, pp. 81–88, 2019, doi: https://doi.org/10.35706/sys.v1i2.2076. [14] g. maulani, d. septiani, and p. n. f. sahara, “rancang bangun sistem informasi inventory fasilitas maintenance pada pt. pln (persero) tangerang,” icit j., vol. 4, no. 2, pp. 156–167, 2018, doi: 10.33050/icit.v4i2.90. [15] b. rumpe, modeling with uml. aachen: springer, 2016. vol. 6, no.1, january 2025 | 10 reinforcement learning-based autonomous soccer agents: a study in multi-agent coordination and strategy development biplov paneru1*, bishwash paneru2, ramhari poudyal3, khem narayan poudyal4 1department of electronics & communication engineering, nepal engineering college pokhara university, nepal 2,4 department of applied science and chemical engineering, tribhuvan university, nepal 3department of information system engineering, purbanchal university, nepal email: biplovp019402@nec.edu.np, rampaneru420@gmail.com, rhpoudyal@gmail.com. khem@ioe.edu.np received: 2024-05-25 | revised: 2024-12-19 | accepted: 2025-01-10 abstract reinforcement learning (rl) approaches, particularly q-learning, have emerged as strong tools for autonomous agent training, allowing agents to acquire optimum decision-making rules through interaction with their surroundings. this research investigates the use of q-learning in the context of training autonomous agents for robotic soccer, a complex and dynamic arena that necessitates strategic planning, coordination, and adaptation. we studied the learning progress and performance of agents taught using q-learning in a series of experiments carried out in a simulated soccer setting. during training, agents interacted with the soccer environment, iteratively changing their q-values in response to observable rewards and behaviors. despite the high-dimensional and stochastic character of the soccer domain, q-learning helped the agents develop excellent tactics and decision-making capabilities. notably, our study found that, on average, the agents required 64 steps to reach a stable policy with an average reward of -1. keywords: q-learning, reward, reinforcement learning, soccer agents i. introduction reinforcement learning (rl) approaches are increasingly being used to teach autonomous entities to complete hard tasks in the fields of artificial intelligence and robotics. one such exciting application is robotic soccer, in which teams of autonomous agents work to achieve predetermined goals, mimicking the dynamics of real-world soccer matches. this study focuses on the creation and analysis of rl-based robotic soccer agents, specifically their capacity to acquire effective tactics for gaming, decision-making, and coordination in a dynamic and competitive setting. this conclusion shows that the agents successfully balanced exploration and exploitation, progressively learning to maximize cumulative rewards while avoiding penalties. furthermore, the observed average reward of -1 implies that, on average, the agents had unsatisfactory results during their learning process, emphasizing the difficulties associated with understanding the complexity of robotic soccer. figure 1. reinforcement learning powered robot involved in soccer p-issn: 2715-2448 | e-issn: 2715-7199 vol.6 no.1 january 2025 buana information technology and computer sciences (bit and cs) mailto:rhpoudyal@gmail.com vol. 6, no.1, january 2025 | 11 soccer players must be able to study and develop basic abilities in order to gain a comprehensive and advanced grasp of the game. these abilities can eventually be combined and utilized to mimic the knowledge of seasoned players. this work explains the application of reinforcement learning, a machine learning approach, to learn the fundamental abilities of intercepting a moving ball. the results of simulation runs on the robocup soccer server were also presented (a. sarje et al, 2004). at the centro universitário da fei, writers were working on a project to compete in the robocup simulation league, which aimed to test reinforcement learning methods in a multiagent domain. the article outlines the squad formed for the robot soccer simulation tournament. they conclude that reinforcement learning techniques are effective in this arena (celiberto et al, 2005). the benefit of rl is the incorporation of a reward system when selecting an action that translates a video frame from a soccer match to one of three potential states. unlike competing techniques, we designed the rl model such that participants' team labels do not need to be explicitly identified. we use a deep recurrent q-network (drqn) to discover the best policy. for effective drqn training, we presented decorrelated experience replay (der), a technique that picks essential events based on the correlations of the experiences recorded in replay memory. experimental findings reveal that computing pass and possession statistics is at least 5.75% and 2.1% more accurate than using similar methodologies (s. sarkar et al, 2023). robots are allocated roles based on the scenario on the gaming field. each job has distinct behaviors and duties. the rl assists the helper and defender in improving their policy choosing abilities during real-time confrontations. the rl system may learn not just how helper assists its colleagues in forming an assault or defense type, but also how to maintain a suitable defensive approach. some trials on the fire simulator and standard platform have shown that the suggested strategy outperforms its rivals (hu c et al, 2020). this study introduces a unique multiagent reinforcement learning (marl) algorithm, nash-learning with regret matching, which uses regret matching to accelerate the well-known marl algorithm nash-learning. it is vital to adopt an appropriate method for action selection in order to balance the relationship between exploration and exploitation and improve the ability of online learning for nash-learning. in a markov game, the combined action of agents using the regret matching method can converge to a set of no-regret points that can be considered as coarse correlated equilibrium, which contains nash equilibrium in essence (y. ma et al, 2009). in comparison to the literature, our research methodology emphasizes a detailed implementation of rl algorithms, particularly in the action selection process. we explicitly outline how actions are chosen based on the current state and exploration strategy, enhancing transparency and reproducibility in algorithmic implementation. this level of clarity ensures consistency and facilitates future research efforts in the field (y. ma et al, 2009). our methodology advances the previous methods with a simpler concept comprising up of a soccer field that was designed with specified dimensions, goals, and boundary conditions to mimic the dynamics of a real soccer game. we employed rl algorithms, namely q-learning, to train autonomous agents to investigate their environment, make strategic decisions, and interact with other agents and the ball. the agents' activities were dictated by the game's current state, with rewards and punishments given in line with established rules and objectives. training iterations were utilized to iteratively alter the agents' rules, resulting in improved performance over time. furthermore, our methodology is consistent with previous research in robotic soccer, leveraging simulation-based methodologies and reinforcement learning (rl) algorithms, notably q-learning, to train autonomous agents. similar to earlier research, we used the pygame package to create a virtual soccer environment, modeling the field with specified dimensions, goals, and boundary conditions to mimic realistic gameplay dynamics. the agents' actions were dictated by the game's present state, with rewards and punishments provided in accordance with established rules and goals. our technique included creating a q-table to hold q-values for state-action pairings, specifying learning parameters including the learning rate (α), discount factor (γ), and exploration rate (ε), and updating q-values with the q-learning update equation. ii. methods to carry out this research, we used a simulation-based technique and the pygame package to develop a virtual soccer environment (f. michaud et al, 1998). the soccer field was modeled with specific dimensions, goals, and boundary conditions to simulate the dynamics of a genuine soccer game. vol. 6, no.1, january 2025 | 12 we used rl algorithms, especially q-learning, to teach autonomous agents to explore their surroundings, make strategic decisions, and interact with other agents and the ball (m. asada, e et al, 1999). the agents' behaviors were dependent on the game's current condition, with rewards and punishments provided in accordance with predetermined rules and objectives. training iterations were used to iteratively adjust the agents' rules and enhance their performance over time. this work uses reinforcement learning (rl) techniques, notably q-learning, to train autonomous agents for robotic soccer games. the simulation system is built on the pygame package, which provides a flexible framework for developing and visualizing soccer situations. 1. action an action made in reaction to the existing situation. in this context, the action might imply whether the left paddle travels up, down, or remains still. the action value is usually expressed as an integer. for example, an action of 0 may correlate to moving the paddle up, 1 to maintaining it stationary, and 2 to pushing it down. 2. reward the reward gained after doing the stated action in the current state. in reinforcement learning, incentives are utilized to communicate the desirableness of a certain state-action combination. a positive reward often denotes a favorable outcome, whereas a negative reward suggests a bad outcome. 3. new state the new state is the consequence of doing the given action in the current state. it represents the status of the environment after the activity has been carried out. the new state is usually decided by the game dynamics and the impact of the action on the surroundings. 4. q-learning q-learning is a model-free reinforcement learning method that determines the best action-selection strategy for a given finite markov decision process (mdp). it learns the importance of doing a certain action in a given condition and strives to maximize the overall reward over time. q-learning is a type of temporal difference learning in which the agent learns from differences between successive estimations of the value function. the equation for q-learning is given by: 𝑸(𝒔, 𝒂) ← 𝑸(𝒔, 𝒂) + 𝜶[𝒓 + 𝜸𝑚𝑎𝑥𝒂𝑸(𝒔′, 𝒂) − 𝑸(𝒔, 𝒂)] (1). where: a) q(s,a) is the estimated value (q-value) of taking action a in state s. b) α (alpha) is the learning rate, controlling how much the q-values are updated after each iteration. c) r is the immediate reward received after taking action a in state s. d) γ (gamma) is the discount factor, representing the importance of future rewards. it determines the balance between immediate and future rewards. e) ′s′ is the next state reached after taking action a in state s. reward: a positive reward is given to the agent when it performs an activity that brings it closer to accomplishing its goal. in the context of a game, this might refer to effectively striking the ball with the paddle, earning points, or stopping the opponent from scoring. positive incentives motivate the agent to carry out similar acts in the future. total reward: sum up all the rewards obtained by the agent. average reward per step: 𝐓𝒐𝒕𝒂𝒍 𝒓𝒆𝒘𝒂𝒓𝒅 / 𝒕𝒉𝒆 𝒕𝒐𝒕𝒂𝒍 𝒏𝒖𝒎𝒃𝒆𝒓 𝒐𝒇 𝒔𝒕𝒆𝒑𝒔 (2). here, 𝑁𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑆𝑡𝑒𝑝𝑠 𝑇𝑎𝑘𝑒𝑛 = 𝑡𝑜𝑡𝑎𝑙 𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑠𝑡𝑒𝑝𝑠. punishment: a punishment, also known as a penalty or negative reward, is provided to the agent when it engages in an activity that pushes it away from its purpose or results in unwanted outcomes. vol. 6, no.1, january 2025 | 13 in the context of a game, this might include missing the ball, allowing the opponent to score, or making ineffective plays. punishments dissuade the agent from engaging in similar behavior in the future. the simulated soccer pitch has predetermined dimensions (600x400 pixels) and includes two paddles representing players, a ball, and boundary lines. the location of each paddle is controlled by a matching agent, while the ball travels dynamically across the surroundings. 5. q-table initialization the q-table stores learnt q-values, which reflect projected future rewards for each state-action pair. initially, the q-table is filled with random values representing all conceivable state-action combinations. the locations of the ball and paddles establish states, whereas actions represent the agents' possible movement possibilities. the training procedure is a continuous game loop in which agents decide actions depending on their current state and learnt q-values. at each iteration, the agent observes the current state, consults the q-table to decide the appropriate action, and then performs the selected action to update the game state. the update_game_state method controls the game's dynamics, such as paddle and ball movement, collision detection, scoring, and resetting when the ball goes out of bounds. this function guarantees that the game proceeds realistically and that the agents receive feedback in the form of incentives depending on their activities. 6. q-value update the q-learning algorithm updates the q-values in the table after each action. the update equation takes into account the observed reward, the highest projected future reward for the next state, as well as the learning rate and discount factor factors. this iterative method allows the agents to learn from experience and eventually improve their decision-making policies. 7. parameter tuning and analysis the performance of rl-based agents is assessed using metrics such as convergence speed, average reward, and gaming efficacy. to enhance learning, characteristics such as the learning rate and discount factor are routinely modified and assessed. furthermore, the effects of explorationexploitation techniques on learning efficiency are investigated. algorithm: a. initialize environment: set up the simulated soccer environment using a suitable library (e.g., pygame) with defined dimensions, player paddles, ball, and boundary lines. b. initialize q-table: create a q-table to store q-values for all state-action pairs. initialize the q-table with random values representing the expected future rewards. c. define learning parameters: 1) set parameters such as the learning rate (α), discount factor (γ), and exploration rate (ε) for the q-learning algorithm. 2) update the q-values in the q-table using the q-learning update equation, incorporating the observed reward and the maximum expected future reward. 3) check termination: check if the game has ended (e.g., ball out of bounds). if so, reset the game to the initial state. d. game loop: 1) start game: begin the game loop to iterate over each time step. 2) observation: obtain the current state of the game, including the positions of the ball and paddles. 3) action selection: based on the current state and exploration strategy (e.g., ε-greedy), select an action using the q-values from the q-table. 4) execute action: move the paddle(s) according to the selected action and update the game state. 5) reward calculation: determine the immediate reward based on the action taken and the resulting state of the game. 8. update q-values a. performance evaluation: 1) monitor convergence: track the convergence of q-values over iterations. 2) evaluate average reward: calculate the average reward obtained per episode or time step. vol. 6, no.1, january 2025 | 14 3) analyze learning progress: assess the effectiveness of the learning algorithm in improving gameplay performance. b. parameter tuning and analysis: conduct experiments to tune the learning parameters (e.g., α, γ, ε) for optimal performance. analyze the impact of parameter variations on learning speed, convergence, and overall gameplay effectiveness. c. iterative training: 1) repeat the game loop for multiple episodes or iterations to allow the agents to learn and refine their strategies over time. 2) continue training until the agents achieve satisfactory performance or convergence. d. results and analysis: 1) evaluate the performance of the rl-based agents based on predefined metrics such as goal scoring rate, defensive efficiency, and overall match outcome. 2) analyze the learned strategies, decision-making processes, and emergent behaviors exhibited by the agents. table 1. learning parameters parameter description value(s) learning rate (α) rate of q-value updates 0.1 discount factor (γ) weight assigned to future rewards 0.9 exploration rate probability of selecting random actions (ε-greedy) 0.1 iii. results and discussions a successful soccer simulation using agents was prepared finally and the reward and other parameters were observed consistently by the authors. this was a kind of successful implementation of rewarding system for the agents and to make them applied with the q-learning method for soccer simulation. figure 2. gaming screen in this work, we looked at the performance of rl-based robotic soccer agents trained with the qlearning algorithm in a simulated environment. the agents' gaming competence and strategic decisionmaking improved significantly during repeated training episodes. the convergence study found that the agents required an average of 64 steps to reach a stable policy with an average reward of -1. this suggests that the agents were successful in developing effective tactics while balancing exploration and exploitation. furthermore, the agents demonstrated adaptive behaviours and trained to maximise cumulative rewards while minimising penalties, demonstrating the effectiveness of q-learning in enabling learning in dynamic and competitive situations. vol. 6, no.1, january 2025 | 15 figure 3. model metrices figure 4. obtained results the plot above depicts three metrics for a reinforcement learning model: total reward, number of steps, and average reward per step. the number of steps taken and total reward appear to be growing simultaneously, whereas the average payout per step fluctuates but remains mainly flat. here are a few explanations of these findings: 1. the model is still learning. when a model is in the early phases of training, the number of steps taken is likely to rise as it explores its surroundings and learns new things. at the same time, the overall reward may rise as the model learns new lucrative activities. the average reward per step may fluctuate as the model makes both excellent and bad judgments, although it is unlikely to change significantly if the model is still exploring extensively. 2. the task is challenging: if the task is difficult or complicated, the model may need to go through several phases to obtain a satisfactory reward. in this situation, both the number of steps done and the overall reward may steadily grow over time as the model's performance increases. 3. the reward function is not pretty well defined and not good enough. if the reward function is not well stated, the model may be unable to learn which activities are actually rewarding. in this situation, the model may go through several stages without making much progress, and the overall reward and average reward per step may not grow considerably. state: the given game log depicts a series of states, actions, awards, and new states experienced during gameplay. let us break down the significance of each component. each state in the log corresponds to the current setting of the gaming environment. it often includes information such as the ball's location, paddle positions, and ball direction. for example, the state (15, 20, 15, 1, -1) might mean that the ball is at position (15, 20), the left paddle is at position 15, the right paddle is at position 15, and the ball is traveling in the positive x-direction (-1) and negative y-direction (-1). figure 5. game log vol. 6, no.1, january 2025 | 16 we can track how the two players' scores change by studying the game log. when the score for the model's side rises, it indicates that the model is being rewarded for effective behaviors. conversely, as the opponent's score rises, it indicates that the model is being penalized for failing to prevent the opponent from scoring. as a result, by examining the score changes in the log, we can determine when the model is being awarded or punished based on its performance in the game. while our technique is comparable to previous studies, it adds new insights by providing a full and detailed explanation of implementation procedures and action decision processes. our study improves comprehension and boosts repeatability of robotic soccer and rl approaches by giving comprehensive explanations and rigorous assessment measures [6]. using simulation-based methodologies and the pygame package, we were able to replicate comparable frameworks used in previous studies [1-4]. these research have shown the effectiveness of rl algorithms, particularly q-learning, in training autonomous agents to navigate dynamic surroundings, make strategic decisions, and interact successfully with other agents and the ball.in summary, our comparative study emphasizes the compatibility of our research methods with current literature, as well as the original additions and advancements presented in our approach. by expanding on known frameworks and addressing critical implementation issues, our study enhances the state-of-the-art in robotic soccer research and underlines the efficacy of rl approaches in autonomous agent training. iv. conclusions to sum up, the game log provides a thorough record of an agent's interactions with its surroundings throughout gaming. the sequence of states, actions, rewards, and new states provides vital insights into the game's dynamics and the implications of various agent actions. analyzing this log can give valuable information for understanding how the game works, such as the movement of game elements like the ball and paddles, the influence of player actions on the game state, and the resultant rewards or punishments. such insights are useful in building and improving reinforcement learning algorithms for training agents to perform effectively in games. reinforcement learning agents can adjust their methods over time by learning from previous experiences logged in the game log, resulting in improved performance and greater points. overall, the game log is an invaluable resource for researching game dynamics, developing efficient reinforcement learning algorithms, and, eventually, improving ai agents' capacities to handle complicated tasks and settings. references a. sarje, a. chawre and s. b. nair, "reinforcement learning of player agents in robocup soccer simulation," fourth international conference on hybrid intelligent systems (his'04), kitakyushu, japan, 2004, pp. 480-481, doi: 10.1109/ichis.2004.81. celiberto, luiz & reinaldo, jr & bianchi, reinaldo. (2005). a reinforcement learning based team for the robocup 2d soccer simulation league. s. sarkar, d. p. mukherjee and a. chakrabarti, "reinforcement learning for pass detection and generation of possession statistics in soccer" in ieee transactions on cognitive and developmental systems, vol. 15, no. 2, pp. 914-924, june 2023, doi: 10.1109/tcds.2022.3194103. hu c, xu m, hwang k-s. an adaptive cooperation with reinforcement learning for robot soccer games. international journal of advanced robotic systems. 2020;17(3). doi:10.1177 /1729881420921324 y. ma, z. cao, x. dong, c. zhou, and m. tan, “a multi-robot coordinated hunting strategy with dynamic alliance,” in proceedings of the chinese control and decision conference (ccdc '09), pp. 2338–2342, chn, june 2009. f. michaud and m. j. matarić, “learning from history for behavior-based mobile robots in nonstationary conditions,” machine learning, vol. 31, no. 1–3, pp. 141–167, 1998. https://doi.org/10.1177/1729881420921324 https://doi.org/10.1177/1729881420921324 vol. 6, no.1, january 2025 | 17 m. asada, e. uchibe, and k. hosoda, “cooperative behavior acquisition for mobile robots in dynamically changing real worlds via vision-based reinforcement learning and development,” artificial intelligence, vol. 110, no. 2, pp. 275–292, 1999. j. hu and m. p. wellman, “nash q-learning for general-sum stochastic games,” journal of machine learning research, vol. 4, no. 6, pp. 1039–1069, 2004. t. fujii, y. arai, h. asama, and i. endo, “multilayered reinforcement learning for complicated collision avoidance problems,” in proceedings of the ieee international conference on robotics and automation, vol. 3, pp. 2186–2191, may 1998. y. wang, cooperative and intelligent control of multi-robot systems using machine learning [thesis], the university of british columbia, 2008. m. wiering, r. sałustowicz, and j. schmidhuber, “reinforcement learning soccer teams with incomplete world models,” autonomous robots, vol. 7, no. 1, pp. 77–88, 1999. m. l. littman, “markov games as a framework for multi-agent reinforcement learning,” in proceedings of the 11th international conference on machine learning, pp. 157–163, 1994. m. j. matarić, “reinforcement learning in the multi-robot domain,” autonomous robots, vol. 4, no. 1, pp. 73–83, 1997. j. h. kim and p. vadakkepat, “multi-agent systems: a survey from the robot-soccer perspective,” intelligent automation and soft computing, vol. 6, no. 1, pp. 3–18, 2000. y. duan, b. x. cui, and x. h. xu, “a multi-agent reinforcement learning approach to robot soccer,” artificial intelligence review, vol. 38, no. 3, pp. 193–211, 2012. r. s. sutton and a. g. barto, reinforcement learning: an introduction, mit press, cambridge, uk, 1998. j. hu and m. p. wellman, “multiagent reinforcement learning: theoretical framework and an algorithm,” in proceedings of the 15th international conference on machine learning, pp. 242–250, 1998. c. j. c. h. watkins and p. dayan, “q-learning,” machine learning, vol. 8, no. 3-4, pp. 279–292, 1992. m. l. littman, “friend-or-foe q-learning in general-sum games,” in proceedings of the 18th international conference on machine learning (icml '01), pp. 322–328, 2001. vol. 5, no.2 june 2024 | 90 implementation of digital invitation applications in the era society 5.0 as a business opportunity micro small to medium (msmes) 1bayu priyatna, 2 agustia hananto, 3 baenil huda 1-3, information systems, faculty of computer science, buana perjuangan university karawang 1 bayu.priyatna@ubpkarawang.ac.id, 2 agustia@ubpkarawang.ac.id, 3 baenil.huda@ubpkarawang.ac.id abstract the rapid development of digital technology in the society 5.0 era and the growth of msmes means that indonesian people have a great opportunity to become msme players, one of which is utilizing digital technology as an innovative product. one of the phenomena that is slowly transforming into a social community is that invitations, originally made using paper, have now become digital invitations. invitations are digital or electronic and can be accessed via computer, tablet, or smartphone devices. digital invitations can be images, videos or text messages created using graphic design or special digital invitation applications. typically, a digital invitation includes information about the event, such as date, time, place and theme. digital invitations can also be decorated with images, icons or illustrations related to the event. the advantages of using a digital invitation include being practical, easy to share, cost-effective, environmentally friendly, and accessible anytime. in contrast to conventional invitations, which are susceptible to damage, require additional fees, limit design choices, and take time to print and send, they generate waste. the research method was descriptive, using a qualitative approach through in-depth interviews with several printed invitation service providers and wedding organizers in karawang. digital invitation application research results can help msmes expand market reach, increase time and cost efficiency, and provide customers with a more interactive and personalized experience. keywords: digital invitation, society 5.0, msmes, e-business. i. introduction in the 21st century, digital technology continues to develop rapidly. smartphones and tablets allow users to access the internet and use applications anywhere and at any time. digital technology has also influenced various aspects of life, including business, education, entertainment and communication. digital technology has also influenced how humans work and socialize in recent years. era society 5.0 creates a human-focused society which can integrate intelligent and innovative technology, such as artificial intelligence (ai), internet of things (iot), robotics, and other digital technologies, to achieve social welfare and prosperity. overall. the society 5.0 concept strengthens collaboration between the public and private sectors. it leads to the development of innovative and holistic technological solutions to address society's complex social and economic problems. indonesia has a lot of potential for creative and innovative businesses, especially in the development of digital technology such as technology startups, e-commerce, gaming, digital marketing, content creation, augmented reality (ar) and virtual reality (vr) and artificial p-issn: 2715-2448 | e-issn: 2715-7199 vol.5 no.2 june 2024 buana information technology and computer sciences (bit and cs) mailto:bayu.priyatna@ubpkarawang.ac.id mailto:agustia@ubpkarawang.ac.id mailto:baenil.huda@ubpkarawang.ac.id vol. 5, no.2 june 2024 | 91 intelligence (ai) technology. this opportunity can certainly be utilized to grow micro, small, and medium enterprises (msmes) to increase the absorption of the workforce. the rapid development of digital technology in the society 5.0 era and the growth of msmes means that indonesian people have a great opportunity to become msme players, one of which is utilizing digital technology as an innovative product. one of the phenomena that is slowly transforming into a social community is that invitations, originally made using paper, have now become digital invitations. a digital invitation is made in digital or electronic form that can be accessed via a computer, tablet, or smartphone. digital invitations can be images, videos or text messages created using graphic design or special digital invitation applications. typically, a digital invitation includes information about the event, such as date, time, place and theme. digital invitations can also be decorated with images, icons or illustrations related to the event. the advantages of using digital invitation include being practical and easy to share, costeffective, more environmentally friendly, and accessible at any time. in contrast to conventional invitations, which are susceptible to damage, require additional fees, limit design choices, and take time to print and send, they generate waste. based on this phenomenon, this research can build a digital invitation application as a manifestation of society 5.0, which can be utilized for business opportunities for the community, improving msmes in indonesia. at the same time, it can also become a forum for transformation for conventional invitation printing businesses. ii. literature review a. information technology information technology is a field that is closely related to technological developments. information technology has positive and negative sides (lawu and ali, 2022). technology can be a tool for improving performance and achieving goals. however, on the other hand, technology can have the opposite effect, so it must be managed wisely (rusdiana and irfan, 2014). according to jogiyanto (wicaksana and saputra, 2021), information technology consists of the words technology and information. technology means applying various equipment or systems to solve problems humans face in everyday life. the word technology is closely related to the term procedure. information is data processed into a more useful form for those who receive it (lawu and ali, 2022). b. digital invitation digital systems refer to information or data stored or transmitted electronically, usually using computer technology. the term "digital" is often used to describe the way computers and other electronic devices process and manipulate information, using binary code (0s and 1s) to represent data (arms 2001). digital technology has changed many aspects of modern life, from communication and entertainment to business and education. digital devices and platforms have enabled new interaction and collaboration and provided access to vast information and resources (foerster-metz et al., 2018). some examples of digital technology include: a. computers, smartphones, tablets and other electronic devices b. internet and world wide web c. social media platforms d. digital media, such as digital photos, videos and music b. e. e-commerce and online shopping a. digital communication tools, such as email, instant messaging, and video conferencing b. online learning platforms and educational resources. digital invitations are invitations made in digital form, either in the form of images, videos or the form online invitations that cannot be physically touched, like conventional invitations that use paper, wood, acrylic and other media (adi ahmad et al. 2022; hadjikhani and lindh 2021). vol. 5, no.2 june 2024 | 92 c. website a website is a collection of various integrated web pages and interrelated files. the web consists of pages or pages, then a collection of pages is called a homepage. the homepage is in the top position, with related pages at the bottom. usually, each page below the homepage is called a child page, which contains hyperlinks to other pages on the web (c p pamungkas, 2015). d. clent-server it is a component that consists of a database application and a dbms server. every activity carried out by the user will be carried out much earlier by the client. next, take action so that the process runs as much as possible yourself. if a process involves data stored in the database, then the client handles the interaction with the server (samsinar and luhur 2018). e. model rpid aplication development (rad) the rapid application development (rad) method, as stated by james martin, consists of four phases: requirements planning phase, user design phase, construction phase, and cutover phase. each phase will be implemented sequentially to develop clis, starting from the requirements planning stage and ending with the cutover phase (kosasi and yuliani 2015). the following is the cycle of rad: fig 1. rapid application development (rad) method (wahyuningrum and pensiunta 2014). the four main phases of rad can be divided into several more specific phases, as depicted in the figure above. the general purpose of phase breakdown is to provide step-by-step information for developers using the rad model to build software. the image shows 2 if-conditional loops; each loop shows how strongly the user is involved in the model. for example, the first loop shows that the requirements planning stage will not advance to the next phase when the information about the system requirements is incomplete and the user decides on the completeness of the information. details about each main phase of rad and the results of each phase will be explained in the next section (wahyuningrum and january 2014). iii. method a. research objects and methodology this research methodology uses a rad system development model approach (rpid application development model); the following is a flowgraph of this research methodology: vol. 5, no.2 june 2024 | 93 fig 2. research methodology b. data collection techniques in collecting data, researchers used several methods, including: a. observation method the observation method is the systematic observation and recording of the symptoms that appear on the research object. this method obtains information by carefully observing and recording the digital invitation. b. interview method what is meant by the interview method is a method of collecting data through observation by conducting verbal questions and answers to business actors, both conventional and digital. c. study of literature the literature study method is to look for data about a thing or variable in scientific journals, books, workshop modules, notes, transcripts, books, newspapers, magazines, minutes, meetings, agendas, etc. c. data analysis data analysis in this research uses quantitative descriptive techniques that describe the monitoring system. data obtained through the instrument was analyzed using quantitative descriptive statistics. this analysis is used to describe the characteristics of the data in each variable. this method makes it easier to understand the data in each process. iv. research results and discussion a. requirements collection the collection of needs in this research was carried out to analyze the phenomenon of changes in the way people view socializing and people's habits in communicating, especially standard patterns in conveying messages officially. the results obtained in this activity are identifying the problems faced problem solving solutions. research methodology data collection technique data analysis research result interview observation study of literature agreement vol. 5, no.2 june 2024 | 94 b. identify the problems faced based on interviews with several parties, it turns out that some people are familiar with the digital invitation system. still, some people and invitation printing business actors do not fully understand the idea of 120 entrepreneurs complaining about increased raw material prices. this long processing time resulted in not receiving all orders: a. increasing raw materials for conventional invitations. b. the manufacturing process takes a long time. c. requires many employees. d. conventional invitation waste is not environmentally friendly. c. problem solving solutions based on the analysis results from several parties, a workflow or criteria for the system to be built can be formulated. the system that will be built is expected to be able to handle problems such as the following: a. cheaper raw materials for invitations result in a decrease in invitation prices. b. the creation process time is much faster just by setting the template c. there is no need for employees to organize invitations because it is user-friendly. d. no waste is generated from digital invitations. in the assembly (creation)/system coding stage, all the objects or materials for the digital invitation application are created. in this research, the programming languages used are php, javascript and html. the following is a display of the program created: a. application front page interface the following is a display image of the front page interface of the digital invitation application: fig 3. application front page interface vol. 5, no.2 june 2024 | 95 b. customer registration page interface the following is a picture of the digital invitation customer registration page interface: fig 4. customer registration page interface c. application login page interface the following is a display image of the digital invitation application login page interface: fig 5. application login page interface d. sample invitation page interface the following is a display image of the digital invitation sample invitation page interface: fig 6. interface example invitation page vol. 5, no.2 june 2024 | 96 e. invitation order page interface the following is a display image of the digital invitation invitation order page interface: fig 7. invitation order page interface fig 8. invoice page interface f. order report page interface the following is a display image of the digital invitation invitation ordering report page interface: fig 9. invitation order report interface vol. 5, no.2 june 2024 | 97 d. testing (evaluation system) this stage is carried out after completing the manufacturing (assembly) stage. hypothesis testing is carried out using structural equation modeling (sem) with the help of the amos version 18 program. this analysis is seen from the significance of the magnitude of the regression weight model and standardized regression weights, presented in table 2: table 2. regression weight estimate s.e c.r p peou  customer/user skills 0.252 0.138 1.790 0.050 peou  resource organization 0.341 0.237 1.422 0.148 peou  display design 0.474 0.171 2.718 0.005 from table 2 above, the results of hypothesis testing can be described as follows: h1: portal design can influence perceived ease of use (perceived ease of use) this hypothesis tests whether the portal design influences perceived ease of use. based on the calculation results in table 2, the significance test for hypothesis 1 is proven to be significant because the probability value obtained is 0.005 or less than 0.05, which means it is important at the 5% significance level with a path coefficient value of 0.252, meaning the relationship between the variables is positive. the quality of the portal design in terms of terminology, interface design and navigation presented by the digital invitation application to its users will influence the perception of ease of use. the digital invitation organization will influence the perception of ease of use (perceived ease of use) this hypothesis aims to test whether the organization of the digital invitation application influences perceived ease of use. based on the calculation results in table 2, the significance test for hypothesis 2 was not proven to be significant because the probability value obtained was 0.148 or greater than 0.05, which means it is not important at the 5% significance level. estimating the organisation's influence on perceived ease of use shows that the path coefficient (standardized regression weight estimate) is 0.341, meaning that the relationship between the variable perceived ease of use and perceived usefulness is negative. this is because system access is easy, fast, and supported by good information resources, making it easier for users to find and obtain varied information. the results of the hypothesis analysis in this research may be due to easy system access. h3: user abilities and skills will influence the perceived ease of use (perceived ease of use) of the digital invitation application this hypothesis aims to test whether user abilities and skills influence perceived ease of use. based on the calculation results in table 2, the significance test for hypothesis 3 was not proven to be significant because the probability value obtained was 0.069 or greater than 0.05, which means it is not important at the 5% significance level. the results of estimating the influence of perceived ease of use on perceived usefulness obtained a path coefficient (standardized regression weight estimate) of 0.252, meaning that the relationship between the user abilities and skills variables on perceived ease of use is negative. based on literature studies, this is caused by the ability and skills of users, in this case, young farmers, who are not good enough, which makes using the digital invitation application not easy and requires time or frequency of use to use the digital invitation application. vol. 5, no.2 june 2024 | 98 v. conclusion training with the digital invitation application in indonesian society has become a focus. it includes technical and practical steps that must be taken to ensure effective and widespread use in the community. digital invitation as their business product. this involves deeply understanding local market preferences and needs and designing a business model accordingly. focus on developing digital invitations with more varied designs and types to create greater attraction. this includes researching local design trends and user preferences and introducing innovative features to meet the diverse needs of the indonesian market. by exploring these aspects, it is hoped that valuable information and solutions can be found that can be applied to increase the acceptance and sustainability of the digital invitation application in indonesian society. reference [1]. prahasta, e. (2001). konsep-konsep dasar sistem informasi geografis. informatika, bandung. [2]. wahyuningrum, t. and januarita, d. (2014) ‘perancangan web e-commerce dengan metode rapid application development (rad) untuk produk unggulan desa’,2014(november), adi ahmad, m.arinal ihsan, hanis dan muharratul. 2022. “view of online digital invitation (an implementation with go-web).” : 52. [3]. arfian, ahmad bayu, ito riris immasari, and asih septia rini. 2022. “perancangan aplikasi undangan digital berbasis website menggunakan codeigniter 4.” jurnal manajamen informatika jayakarta 2(1): 1. [4]. arms, william y. 2001. digital libraries. mit press. [5]. fachri, barany, and risky wahyu surbakti. 2021. “perancangan sistem dan desain undangan digital menggunakan metode waterfall berbasis website (studi kasus: asco jaya).” journal of science and social research 4(3): 263. [6]. foerster-metz, ulrike stefanie et al. 2018. “digital transformation and its implications on organizational behavior.” journal of eu research in business 2018(s 3). [7]. hadjikhani, annoch isa, and cecilia lindh. 2021. “digital love – inviting doubt into the relationship: the duality of digitalization effects on business relationships.” journal of business & industrial marketing 36(10): 1729–39. https://doi.org/10.1108/jbim-05-20200227. [8]. immasari, i. r., & arfian, a. b. 2022. “rancang bangun aplikasi undangan digital pernikahan dengan menggunakan codeigniter.” journal of information system, applied, management, accounting and research, 6(3), 521–531 6(3): 521–31. http://journal.stmikjayakarta.ac.id/index.php/jisamar/article/view/518. [9]. kosasi, sandy, and i dewa ayu eka yuliani. 2015. “simetris : jurnal teknik mesin, elektro dan ilmu komputer.” simetris: jurnal teknik mesin, elektro dan ilmu komputer 6(1): 27– 36. [10]. lawu, suparman hi, and hapzi ali. 2022. “perencanaan strategis sistem informasi dan teknologi informasi dengan pendekatan model: enterprice architecture, ward and peppard.” indonesian journal computer science 1(1): 53–60. [11]. samsinar, samsinar, and universitas budi luhur. 2018. “barang dengan metodologi berorientasi obyek studi kasus : pada pt . moiko tasindo.” (march). [12]. sri asfirawati halik. 2022. “pelatihan membuat undangan digital di era covid-19 training on making digital invitations at covid-19 era.” abdimas galuh 4: 537–42. [13]. wahyuningrum, tenia, and dwi januarita. 2014. “perancangan web e-commerce dengan metode rapid application development ( rad ) untuk produk unggulan desa.” 2014(november): 81–88. p-issn : 2715-2448 | e-issn : 2715-7199 vol.4 no.1 january 2023 buana information technology and computer sciences (bit and cs) 11 | vol.4 no.1, january 2023 an enhanced bio-inspired aco model for fault-tolerant networks samuel w lusweti 1 masinde muliro university of science and technology email: lusweti015@gmail.com collins o odoyo 2 masinde muliro university of science and technology email: codoyo@mmust.ac.ke ‹β› dorothy a rambim 3 masinde muliro university of science and technology email: drambim@mmust.ac.ke abstract—this research mainly aimed at establishing the current functionality of computer network systems, evaluating the causes of network faults, and developing an enhanced model based on the existing aco model to help solve these network issues. the new model developed suggests ways of solving packet looping and traffic problems in common networks that use standard switches. the researcher used simulation as a method of carrying out this research whereby an enhanced algorithm was developed and used to monitor and control the flow of packets over the computer network. the researcher used an experimental research design that involved the development of a computer model and collecting data from the model. the traffic of packets was monitored by the cisco packet tracer tool in which a network of four computers was created and used to simulate a real network system. data collected from the simulated network was analyzed using the ping tool, observation of the movement of packets in the network and message delivery status displayed by the cisco packet tracer. in the experiment, a control was used to show the behavior of the network in ideal conditions without varying any parameters. here, all the packets sent were completely and correctly received. secondly, when a loop was introduced in the network it was found that the network was adversely affected because for all packets sent by the computers on the network, none of them was delivered due to stagnation of packets. in the third experiment, still, with the loops on, a new aco model was introduced in the cisco packet tracer used to simulate the network. in this experiment, all the packets sent were completely and correctly delivered just like in the control experiment. keywords—aco, packets ,loops, networks, algorithm abstrak—penelitian ini bertujuan untuk menetapkan fungsionalitas sistem jaringan komputer saat ini, mengevaluasi penyebab kesalahan jaringan, dan mengembangkan model yang ditingkatkan berdasarkan model aco yang ada untuk membantu memecahkan masalah jaringan ini. model baru yang dikembangkan menunjukkan cara-cara untuk memecahkan perulangan paket dan masalah lalu lintas di jaringan umum yang menggunakan standard switches. peneliti menggunakan simulasi sebagai metode untuk melakukan penelitian ini dimana algoritma yang disempurnakan dikembangkan dan digunakan untuk memantau dan mengontrol aliran paket melalui jaringan komputer. peneliti menggunakan desain penelitian eksperimental yang melibatkan pengembangan model komputer dan mengumpulkan data dari model tersebut. lalu lintas paket dipantau oleh alat cisco packet tracer di mana jaringan empat komputer dibuat dan digunakan untuk mensimulasikan sistem jaringan nyata. data yang dikumpulkan dari simulasi jaringan dianalisis menggunakan alat ping, pengamatan pergerakan paket dalam jaringan dan status pengiriman pesan yang ditampilkan oleh cisco packet tracer. dalam percobaan, kontrol digunakan untuk menunjukkan perilaku jaringan dalam kondisi ideal tanpa memvariasikan parameter apa pun. semua paket yang dikirim diterima dengan lengkap dan benar. ketika sebuah loop diperkenalkan pada jaringan, ditemukan bahwa jaringan terpengaruh secara negatif karena untuk semua paket yang dikirim oleh komputer di jaringan, tidak ada satupun yang dikirim karena stagnasi paket. dalam percobaan ketiga, masih dengan loop aktif, model aco baru diperkenalkan di pelacak paket cisco yang digunakan untuk mensimulasikan jaringan. dalam percobaan ini, semua paket yang dikirim dikirim dengan lengkap dan benar seperti pada percobaan kontrol. kata kunci—aco, packets ,loops, networks, algorithm i. introduction one can rarely do anything with data that doesn’t involve a computer network since these networks enable us to share information and other resources [1]. however, these networks have to be maintained well to keep providing users with these functionalities. these computer networks can fail to work especially when a network device fails or the communication link malfunctions or is being overused against its capacity [2]. the existing networks are more complex than the conventional networks therefore it becomes hard to create, install, manage and keep them up and running efficiently all the time [3]. due to these challenges, new technologies are needed to be employed in managing these dynamic networks that are at the center of business transactions today. there exist similar problems in real-life situations and their biological solutions which are naturally evolved and can also be applied in networking paradigms to help curb the drawbacks [3]. these are commonly referred to as bio-inspired systems. a bio-inspired system depicts a strong relationship between a proposed algorithm aimed at solving a certain problem and a biological or natural system possessing similar capabilities [4]. there exists a necessity to employ bio-inspired systems in computer networks because 12 | vol.4 no.1, january 2023 living organisms like the ant colonies look better organized in their daily activities than the current internet [4]. this is because of the resilience to failure by biological systems to both internal and external environmental factors [5]. there exist many optimization algorithms including particle swarm optimization, genetic algorithm, leaping frog among others. in general, aco is the most popular and most successful algorithm that has ever been used in combinatorial optimization problems [6]. ii. literature review 2.1. loops loops can occur in a network whenever there exists more than one path (redundant links) at layer 2 between two endpoints or there are multiple connections between any two switches in the network or there is a physical connection between two active ports of the same switch. this loop creates broadcast storms during the forwarding of broadcasts and multicasts by switches. these switches repeatedly rebroadcast the message hence flooding the network [7]. this means that packets that are sent along a given path will eventually be stuck in that network forever making cycles, [8] unless some mechanisms are invoked which will help in flushing such packets out of the cycle. whenever you are using link-state protocols for instance ospf, the forwarding loops may transiently occur if the routers in use adapt their own forwarding route tables in response to a change in topology [9]. the adverse effects of the loop on ethernet switched networks include a reduction in bandwidth, memory clogging, and packet loss [7]. these and many other challenges in computer networks need to addressed using more intelligent mechanisms especially the use of metaheuristics like bio-inspired systems. this research paper concentrates on packet loops and provides an enhanced bioinspired mechanism to help solve the problem. 2.2 bio-inspired systems biology is frequently being employed as an inspiration for research in computer science [2] and other fields like engineering, mathematics, energy [10], and business. ant algorithms are in use today mainly to solve optimization problems instead of the problems and challenges that they were originally or initially developed to solve [11]. a direct approach to getting a solution to the combinatorial optimization problems is an exhaustive search [12], whereby the agents in these biological systems whose analogy is used to create bio-inspired systems, enumerate all possible solutions and choose the best one. the method of ants which is a bio-inspired system has proved to outshine other generalpurpose algorithms for optimization like the genetic algorithms. this is evident especially when employed in combinatorial optimization problems which require the interaction of cooperating agents [11]. a good number of such intelligent algorithms are therefore being developed with the aim of solving various complex problems. whereas some studies try to explore the application of bio-inspired algorithms theoretically, others are continuously working to improve the functionality of the algorithms [10]. this becomes a green light for the growth and development of artificial intelligence systems which mimic the behavior of ants, especially during foraging. these algorithms include neural networks, particle swarm, genetic algorithm, and ant colony optimization algorithm among others [10]. 2.2.1 aco architecture and design during the early years of 1990s, there was an introduction of the aco model proposed by m. dorigo and his companions as a metaheuristic optimization algorithm that is naturally inspired for solving hard combinatorial optimization problems [13]. aco is in the category of metaheuristics which are probabilistic algorithms for obtaining good solutions to hard combinatorial optimization problems with a reasonable time of computation [14]. the other examples of metaheuristics include tabu-search, evolutionary computation, and simulated annealing [15] [16]. the foraging behavior of real ants inspired the creation and deployment of aco in various fields. aco is a very popular and modern optimization paradigm that gets its motivation from the scenario of how ant colonies find the shortest routes between their home and the food source [6]. during the process of searching for food (foraging), the forward ants start by randomly exploring the environment that surrounds their nest before they locate the source of food and deposit pheromone trails [17]. despite its well-known popularity and application, the theory of aco is still under development and is in its infancy stages; therefore a solid foundation of its theory is required [18]. the aco model has been modified in many forms since its introduction. however, these forms have a common architecture. the following figures show an image of real ants in a colony foraging and a flow chart of activities carried out by ants in an aco algorithm during their foraging activities. 2.2.2 how aco works ants are self-organized biological systems exhibiting three main principles of self-organization mechanisms which include interaction between individuals, feedback loops, and local state evaluation [19]. aco algorithm works following an indirect communication among the simple agents of a colony, known as artificial ants, enabled by their artificial pheromone trails as their media of communication [20]. the foraging ants thus communicate by laying pheromone chemicals on the ground as they search the environment for food [18]. the other ants are consequently attracted by the laid pheromone trails and therefore tend to closely follow previous ants. in the case whereby the foraging ants discover different routes between the nest and a source of food, the shortest path typically gets filled with pheromone quicker than the longer path [18]. a fascinating aspect of ants is their ability to find the shortest routes to the source of food. this is made possible only by the ability of these ants to follow the laid down pheromone trails by the predecessor ants taking in mind that most ants are almost blind meaning visual aspects are not in use [21]. the more the ants take the shortest route, the more pheromone gets deposited, until almost all ants follow the shortest path [18]. 2.3 application of aco in networks there have been many areas in which bio-inspired systems have been successfully applied, some of which have been mentioned in the sections above. however, a few of them have been applied in computer networks to enhance their functionality and improve their resilience to faults. the main area of concern being adaptive routing. a good example is the aco algorithm for network analysis and adaptive routing 13 | vol.4 no.1, january 2023 [10]. many issues in networking are formulated as multidimensional optimization problems. as dimensions of networks are increased both in terms of the number of nodes and spatially, the centralized control of communication in these networks becomes impractical. in comparison with biological systems, an individual alone is of less interest compared to the collective behaviour of the system of a larger number of the same individuals like the ant colonies [10]. as a result, bio-inspired systems like aco have been developed and successfully applied in network node research and design due to the appealing analogies existing between biological systems and large computer networks [2]. aco has been applied in computer network routing problems in different formats such as antnet, anthocnet, acr [6], and antbasedcontrol(abc) [22] among other algorithms. abc was the first routing algorithm applied in circuit-switched networks for instance telephone wire networks [23]. this algorithm was deployed on a simulated version of the british telecom network, which formed the basis [20] of more research on this area of network routing problems. a highly successful application of the aco to the dynamic routing problems is the antnet algorithm, which was proposed by marco d and di caro [24] [25] [26]. this antnet being successful was recommended and applied in packet-switched networks in this case the internet for adaptive routing [20]. later on, the antnet became widely useful in mobile adhoc networks, to solve the routing problems and was used as anthocnet posting exemplary results [27]. antnet [28] and anthocnet [29] are two today very well-known aco-based routing algorithms. antnet is a proactive routing algorithm while the anthocnet is a reactive routing algorithm. they possess a very high delivery rate and find paths whose lengths are very close to the length of the shortest route [28],[30]. 2.4 challenges faced by aco if aco is used in network routing, the packets mimic the real ants. however, it still has some challenges. in this case, a problem arises when an ant (packet) is stuck in a cycle (loop) and is forced to revisit an already visited node. these loops bring about the stagnation of ants [31] and are undesirable in networks because they cause packet latencies [6] and the packets caught in a loop are eventually destroyed when aco is implemented. the large number of routes or paths makes it complex to manage the routing tables while concurrently increasing the probability of having packet loops [32]. although research was done and a new aco was developed known as maco, the research does not claim that maco will fully eradicate stagnation in ants, thus offering a basis for more research to be done on aco [31]. these reasons necessitated research on how best to re-route the packets that are caught into a loop and stagnated without destroying them and by preventing stagnation and still making them reach their intended destination in the new aco algorithm. iii. method a model was developed using python and pygame simulator to show the movement of ants. this study used the agile family as a model of sdlc since these methods are meant to quickly adapt to changing requirements, and minimize costs of production while upholding the quality of the software under development [33]. agile is a combination of both incremental and iterative types of sdlc [34]. iv. results and discussion in this chapter, the researcher presents the results of the simulation experiments performed. in these experiments, the researcher used the cisco packet tracer simulator to test various configurations of the network as shown below. first of all, the packet tracer is launched and a simple network configuration of four computer devices (two desktop computers and two laptops) are configured in the network. a router is used to bridge between two different network classes each having two computers. these classes include b having a default gateway of 172.16.0.1 and class c with a default gateway of 192.168.0.1. these computers are interconnected on the network using two 24-port. the so-developed network of computers is then subjected to different conditions and tested as shown by the screenshots below. 4.1. simple-aco the old model called simple-aco was applied in various fields including computer networks. s-aco ants were used to implement loop elimination which in turn improve the performance of the system [12]. while moving backwards through the path that in in their memory, the ants update the pheromone concentration on the paths they pass through. the following example shows how ants in s-aco eliminate loops as using backward pass presented by marco dorigo [12]. figure 1. final route, no loop but also node 4 is removed in aco, the private memory of ant is used to guarantee the probability of an ant building a feasible solution. however, during the process of finding these feasible routes, loops may appear making the ant rotate in a cycle for long. as a result, the ant ends up wasting time and resources [6]. when trying to avoid these loops, if an ant is forced to go back to an already toured node, the nodes having that cycle are removed from the private memory of the ant and information about the nodes is completely destroyed. if an ant moves in a loop for a time which is greater than half its age (time to live), the ant is consequently destroyed [6]. the model describes the movement of an ant from source to destination. from the flow diagram, if an ant is forced to return to an already visisted node, the node or ant will be destroyed. this saves the system from staying in a loop forever, but the killed ant never reached its destination. if the node is detryoed, then it is removed from the network making the node lose network connection. 14 | vol.4 no.1, january 2023 figure 2. flow chart of original aco model showing how to eliminate loops pseudocode if (k ∈ v )/∗ check if the ant is in a loop and remove it ∗/ hops_cycle ← get cycle_ length(k,v); hops_fw ← hops_fw – hops_cycle; else hops_fw ← hops_fw + 1; v [hops_fw] ← k; t ,[hops_fw] ← tk→n; end if 4.2. enhanced aco model showing physical movement of ants figure 3. loops formed are assumed to be non-existent (for instance node 4 is not removed) in the figure above, if a loop is formed in the modified aco model, it will be assumed by the ants as the ants will still be reroute out of the loop if they are caught in the loop for more than half of their ttl. in this case no ant will be killed and no node will be destroyed. the ants are made to be a little more intelligent in that, when an ant is caught in a cycle, its assigned time to live will be used to determine if the ant is in a loop. if the ant takes rotates in a loop for a time interval which is greater than half of its ttl, then it is rerouted out of the loop to its intended destination. figure 4. flow chart enhanced aco model showing how to eliminate loops pseudo-code of enhanced aco model if (k ∈ v )/∗ check if the ant is in a loop and remove it ∗/ hops_cycle ← getcycle length(k,v; hops_fw>ttl/2; else hops_fw ← hops_fw + 1; v [hops_fw] ← k; t ,[hops_fw] ← tk→n; end if in the enhanced algorithm above, if an ant is caught in a cycle, hops_fw>ttl/2 , the else statement function hops_fw ← hops_fw + 1 is invoked to get the ant out of the loop unconditionally. the ant selects the next node described in its foraging instructions. this enhanced algorithm is then applied in the simulated computer networks to help reroute packs intentionally forced to return to selected nodes. this forced return is done by introduction of forced loops. 4.0 control experiment pinging the devices in the control experiment for instance from laptop 1 having ip address 172.16.0.3 to laptop 0 and pc 0 having ip addresses 192.168.0.3 and 192.168.0.2 respectively, still shows communication taking place as shown by the command prompt screen in figure 5 below. 15 | vol.4 no.1, january 2023 figure 5. control results 4.1 simulator with loops but without the algorithm the simulated network is configured as shown below. there is a modification from the control experiment by the introduction of physical loops on the network. figure 6 below shows the configuration. after simulating packets in this network, the message remained in progress forever as shown in the bottom right corner of the figure below. secondly, the messages in the network simulator kept on rotating between the loops created and the two switches 0 and 4. figure 6. loops without aco result 1 when we open the command prompt and ping from laptop 1 having ip address 172.16.0.3 to laptop 0 and pc 0 having ip addresses 192.168.0.3 and 192.168.0.2 respectively, the icmp message fails to be delivered and shows time out on the command prompt and failed on packet tracer as shown in figure 7 below. figure 7. loops without aco result 2 4.2 simulator with loops and with the algorithm the researcher was able to run the algorithm and then simulated the network having the loops. secondly, the main algorithm is imported in the programming mode under the tcp python package of each computer device running in the networking. the algorithm displays the activity of random movement of ants on the pygame. the simulation mode of the cisco packet tracer shows message delivery in the network with green ticks as shown in the figure below. at the bottom right corner of the simulator, there is an indication of the successful delivery of messages. figure 8. loops with aco result 1 pinging the devices for instance from laptop 1 having ip address 172.16.0.3 to laptop 0 and pc 0 having ip addresses 192.168.0.3 and 192.168.0.2 respectively, indicates communication taking place as shown by the command prompt screen below. figure 9. loops with aco result 2 4.4 interpretation of results in the above simulation, the researcher used three experiments to test and validate the data collected. the first which is a control experiment provides an ideal network with no variations in the network working under optimal conditions. it was discovered that the network messages were delivered properly between the devices on the network without any problem showing 100% success for the 4 messages sent over the network. in the second experiment, the researcher subjected the simulated network to forced loops as it normally happens in colleges and universities (mostly done by students). in this experiment, it was discovered that the network went completely down and could 16 | vol.4 no.1, january 2023 not recover from the looping packets which completely failed to be delivered to the intended destination. pinging the network showed that the 4 messages send had a 100% failure rate. in the third experiment, the aco algorithm was simulated alongside the network with the loops still on. this experiment showed packets being delivered normally in the simulation mode by showing success to all icmp packets sent. the ping tool also showed a 100% success rate for all the 4 test packets sent by various devices on the network. this shows that the algorithm intervened to reroute the packets that would have otherwise gotten locked up in the loop as shown in experiment 2 where the packets kept rotating in the loop thereby breaking communication between the computers on the network. the algorithm helped the network in rerouting any packets that would be caught in the loop to reach their destination. from the above experiments, five average round trip times(rtt) were collected from control experiment and five from enhanced experimen. the control experiment had the following rtt as shown in figure 8 and figure 9 (6,0,0,2,0) averaging to 8/5 =1.6ms. the experiment of enhanced algorithm showed the following results as rtt as witnessed from figure 16 and figure 17 (0,0,2,0,16) avearaging to 18/5=3.6ms. although the algorithm’s results show that the packets get to destination on average 2 ms later than in the case of the control, it is still witin the allowed rtt. the 2ms delay may have been brought about by the cycling packets before getting out of the loop. for optimal network performance , studies have shown that rtt<500ms [75]. 4.5 how the enhanced aco algorithm works the aco model was first developed and applied in computer networks to help the network packets reach their destination very fast using the shortest route possible by mimicking the ants through their foraging behavior and intelligence. however, the model could not solve the problem of stagnation of ants and thus the new model developed in this research helps to get the stagnated ants out of the loop. the algorithm is then imported into the cisco packet tracer and the packet tracer translates the ants into real packets of the network. when the packets are caught in a loop, they use their random intelligent movement to get out of the loop, once they realize there are talking too long to reach their destination. the movement of the packets (similar to the ants after importing into packet tracer) is as shown below. figure 10. ants in model foraging figure 10 above shows packets finding the routes from the source of food to the nest (dest). figure 11. ants in model caught in a loop the figure above shows packets (ants) caught in a loop for sometime. figure 12. ants getting out of the loop this figure shows (packets) ants getting out of the loop after rotating in the loop for a few moments. the will be unnoticeable in the real network since it happens at a very high speed. 4.6 discussion according to research done by kwang hong sim and weng hong sun [31], when implementing aco the following issues arise: stagnation of ants and poor adaptiveness of ants. stagnation will occur if a network reaches its convergence too early (prematurely). this leads to the following problems: 1) congestion of packets (ants) on the optimal path, 2) reduction of probability of ants choosing other routes. the research further indicated that congestion and low probability of choosing other paths by ants will consequently lead to the following problems in dynamic computer networks: a) the selected path may become non-optimal if congested b) the congested path may be disconnected due to network failure. due to these issues, sim and sun noted that the network would eventually have a low degree of sensitivity due to changes in topology or link failure. in their research, sim and sun [31] proposed a new improved model called maco which they noted that it enhanced the adaptiveness of the network and reduced the chances of stagnation of ants. in maco, the ants are made to deposit pheromones of different colors and observed that ants will follow routes with similar colors reducing congestion. however, they conclude that their research does not guarantee that maco would fully eradicate stagnation of ants as it only reduced the chances of stagnation. secondly, another research was done by gianni di caro on the application of aco to adaptive routing in telecommunications networks [8]. the research showed that loops that lead to stagnation of ants should be avoided as much as possible because they cause very high packet 17 | vol.4 no.1, january 2023 latencies. to implement loop avoidance, di caro explained using an equation for avoiding loops in networks. for instance, if an ant is forced to return to a node that it had previously visited, the nodes that composed the loop are taken out of the memory of the ant, and all information describing those nodes is destroyed. consequently, if an ant moves on a cycle for a time that is greater than its half-life, the ant is destroyed. this destruction solves the problem of loops; although, another problem arises whereby the affected nodes and ants(packets) will be destroyed which means they will not communicate over the network as intended. considering the finding of the above two kinds of research, the researcher carried out this study to help strike a balance between stagnation and the destruction of ants. from the results above, the problem of stagnation of ants is more highly reduced than that of maco suggested by sim and sun [31] by making sure the ants that are caught in a loop simply get out of the loop using a new optimal route as shown in figure 10. this will help us make sure that all the packets in the network reach their intended destination as opposed to destroying stagnated ants as found out in the research by di caro [8]. v. conclusion and recommendations the results show that this research can add some contribution to the body of knowledge. when well adopted in schools, colleges, homes, and offices, the algorithm produced in this research study can go a long way in improving the reliability of the computer networks that are constantly under attack by forced loops. however, the algorithm does not completely guarantee that it will solve all the problems faced by computer networks. further research needs to be done to improve it further. references [1] meyers, m., "introducing basic network concepts." 2010, https://www3.nd.edu/~cpoellab/teaching/cse40814_fall 14/networks.pdf (accessed june. 3, 2022). [2] hylsberg r. jacobsen, q. zhang,t. skjødeberg toftegaard, "bioinspired principles for large-scale networked sensor systems: an overview", sensors, vol. 11, no. 4, pp. 4137-4151, 2011. doi: 10.3390/s110404137. [3] jisha mrriam jose, "bio-inspired networking," 2011 hyperlink "http://dspace.cusat.ac.in/jspui/bitstream/12345678 9/3195/1/bio-inspired%20networking.pdf" http://dspace.cusat.ac.in/jspui/bitstream/123456789/ 3195/1/bio-inspired%20networking.pdf (accessed jan. 23, 2020) [4] jamalipour a. "bio-inspired networking", the university of sydney: ieee, 2009. hyperlink "http://www.it.is.tohoku.ac.jp/~kato/workshop200 9/01.pdf" http://www.it.is.tohoku.ac.jp/~kato/workshop2009 /01.pdf (accessed feb. 20, 2020) [5] dressler, i.f., "self-organization in autonomous sensor/actuator networks [selforg." 2010. [6] caro d.g., dorigo m. "ant colony optimization and its application to adaptive routing in telecommunication networks, 2004". (doctoral dissertation, phd thesis, faculté des sciences appliquées, université libre de bruxelles, brussels, belgium). [7] iufoaroh s.u, chukwumaobi .o,oranugo c.o "tracking the effects of loops in a switched network using rapid spanning tree network," international journal of research in electronics and communication technology (ijrect 2015), vol. 2, no. 3, 2015. [8] balakrishnan h, "network routing ii failures, recovery, and change," 2009. hyperlink "http://web.mit.edu/6.02/www/s2009/handouts/l21 .pdf(accessed" http://web.mit.edu/6.02/www/s2009/handouts/l21. pdf(accessed jan. 6, 2021) [9] francois p., bonaventure o. "avoiding transient loops during the convergence of link-state routing protocols." ieee/acm transactions on networking, 15(6), 2007, pp. 1280-1292 [10] kumar k.a, "bio inspired computing – a review of algorithms and scope of applications", expert systems with applications, vol. 59, pp. 20-32, 2016. doi: 10.1016/j.eswa.2016.04.018. [11] afshar m.h, "a new transition rule for ant colony optimization algorithms: application to pipe network optimization problems", engineering optimization, vol. 37, no. 5, pp. 525-540, 2005. doi: 10.1080/03052150500100312. [12] marco d., thomas s., "ant colony optimization", london: bradford book, 2004. [13] dorigo, m., vittorio m.,alberto c., "ant system: optimization by a colony of cooperating agents." ieee transactions on systems, man, and cybernetics, part b (cybernetics) 26, no. 1 (1996): 29-41 [14] glover f., gary a., "handbook of metaheuristics." vol. 57. springer science & business media, 2006 [15] glover f, "tabu search—part ii,," orsaj. comput., vol. 2, no. 1, pp. 4-32, 1990. [16] glover, f., laguna, m.,"tabu search kluwer academic". boston, texas, dordrecht, 1997. [17] dorigo m., blum c., "ant colony optimization theory: a survey," theoretical computer science, vol. 344, pp. 243-278, 2005. [18] sudholt, d. and thyssen, c., "running time analysis of ant colony optimization for shortest path problems". journal of discrete algorithms, 10, 2012, pp.165-180 [19] dressler, f., "benefits of bio-inspired technologies for networked embedded systems: an overview". in dagstuhl seminar proceedings. schloss dagstuhlleibniz-zentrum fr informatik. 2006. [20] dorigo, m., stützle, t.,. "ant colony optimization: overview and recent advances. handbook of metaheuristics", 2019, pp.311-351. [21] cordon o., herrera f., stutzle t. "a review of ant colony optimization metaheuristic: basis, models and new trends," 2002. 18 | vol.4 no.1, january 2023 [22] mandeep kaur, a.s "a review on ant colony optimization in manet," international journal for science and emerging, vol. 19, no. 1, pp. 1-6, 2014. [23] schoonderwoerd, r., holland, o.e., bruten, j., rothkrantz,. "ant-based load balancing in telecommunications networks". adaptive behavior, 5(2), 1997, pp.169-207. [24] marco. d., gianni c., "ant colonies for adaptive routing in packet-switched communications," in fifth international conference on parallel problem solving from nature, 1998. [25] marco. d., gianni c., "mobile agents for adaptive routing.," in proceedings of the 31st international conference on system sciences, 1998. [26] marco. d., gianni c., " antnet: distributed stigmergic control for communication networks," journal of artificial intelligence research, vol. 9, pp. 317-365, 1998 [27] gianni d., frederick d. ,luca m.g , "using ant agents to combine reactive and proactive strategies for routing in mobile ad hoc networks," international journal of computational intelligence and applications, vol. 5, no. 2, p. 169–184, 2005. doi/10.1142/s1469026805001556. [28] marco. d., gianni c., "ant colonies for adaptive routing in packet-switched communications networks.," in proceedings of the 5th acm international conference on parallel problem solving from nature, p. 673–682, 1998. [29] di caro, g., ducatelle, f., gambardella, l.m. "anthocnet: an adaptive nature‐inspired algorithm for routing in mobile ad hoc networks". european transactions on telecommunications, 16(5), 2005, pp.443-455. [30] gunes m., "ara the ant-colony based routing algorithm for manets.," in proceedings of the 2002 international conference on parallel processing workshops, pages, pp. 79-89, 2002. [31] sim, k.m., sun, w.h., november. "multiple ant-colony optimization for network routing". in first international symposium on cyber worlds, 2002. proceedings, 2002, (pp. 277-281). ieee. [32] maniezzov., "ant colony optimization", 2001. [33] young d., "sofware development methodologies," 2013. [34] sharma, s., sarkar, d., gupta, d.,." agile processes and methodologies: a conceptual study". international journal on computer science and engineering, 4(5), 2012, p.892. [35] rasmussen j., "what is the basis for classifying a call as poor in lync 2013 qoe?," 20 september 2013. [online]. available: https://docs.microsoft.com/enus/archive/blogs/jenstr/what-is-the-basis-forclassifying-a-call-as-poor-in-lync-2013-qoe. (accessed july. 2, 2022). vol. 6, no.1, january 2025 | 46 p-issn: 2715-2448 | e-issn: 2715-7199 vol.6 no.1, january 2025 buana information technology and computer sciences (bit and cs) a novel multi-level perceptron for accurate heart stroke diagnosis muhammad khubaib1, muhammad zaman2*, tanzila kahkashan3, anam zahoor4, narges shahbaz5, shahzad shoukat6, fahma nisar7 1,2,4 faculty of computer science, university of lahore, 10 km lahoresargodha rd, sargodha, pakistan. 3 faculty of computing, universiti teknologi malaysia, johor bahru 81310, malaysia. 5 department of education, university of education, lahore, pakistan. 6,7 department of computer science, comsats university islamabad, islamabad, pakistan. e-mail: mkhubaib141@gmail.com, muhammad.zaman@cs.uol.edu.pk*, tanzila.kehkashan@cs.uol.edu.pk, tanzila@graduate.utm.my, malikanam06@gmail.com, nargesshahbaz20137@gmail.com, karishmaaslam078@gmail.com, fahmanisar001@gmail.com received: 2024-06-23 | revised: 2024-12-10 | accepted: 2025-01-30 abstract heart is a very important part of human body, it supply blood to body. if the heart fail down the person cannot survive. this is very important to diagnose the heart disease timely to start proper treatment. to dingoes this disease manually takes time a lot and budget of the patient. traditionally the patient have to go through form different test then he have to give medical history to the doctor then the doctor make decision about their disease and then the treatment start. in the developing countries especially like pakistan the income of the people are too much low and they cannot offered different type of expensive tests like ecg etc. in this way the disease cannot detect timely and cannot treated properly. the heart stroke can be predicted by analyzing different attributes like blood pressure, cholesterol age etc., this is a best and easy way to predict heat stroke timely. different types of machine learning and deep learning algorithms are used for heart stroke predictions. in this paper we purposed novel multilayer-perceptron (mlp) that are efficient in classification and in heart stoke prediction that model achieve high accuracy of 99% keywords: heart disease detection, cardiovascular disease, classification algorithms, medical imaging i. introduction heart disease is fatal diseases across the world that cause a large number of deaths [1]. this is medical condition when heart stop working properly and cannot pump the proper amount of blood to the body parts [2]. the primary symptoms of this disease is chest pain, irregular heartbeat, swollen feet [3]. the current techniques are not much efficient to detect heart disease early so researcher try to find new technologies to overcome these issues and want to detect disease more accurately and timely [4]. this is a big issue of diagnosing the disease timely while using current techniques when the medical expert not available at that time [5]. if the disease can detect timely this is a big chance to treated this and can save the life of patient [6]. there are millions of people that are cause of heart disease are diagnosed yearly [3]. there is also a big strength of people that are stroked by this fatal disease in united states [1]. the tradition diagnosis of heart disease includes collecting the medical history of patient, report of physical health and then all these are analyzed by the medial specialist, but this is too much expensive and time taking process [1]. the patient has to go laboratories to conducting many tests to finding physical reports, this is also a vol. 6, no.1, january 2025 | 47 type of burden in form of money to the patient [7]. the heart disease identification is complex task due to tests and other details [8]. in developing countries this is also a big issue for patients to get affordable diagnoses due to low income and other financial issues, the health care equipment’s and other health facilities are also limited in these countries [9]. there are four major methods that are used now a days for heart disease diagnosis which include ecg, test of exercise stress, x rays and coronary angiograms [10]. by analyzing different attributes now it is possible to predict the heart store timely these attributes include blood pressure , cholesterol level, age , gender , smoking routine etc. the traditional method for attribute analysis of a patient is time consuming and difficult. now a day’s artificial intelligence play an important role in health different machine learning and deep learning algorithms and technologies are now being used to overcome the issues of previous manual methods. in heart stroke prediction it play an important role. by using different machine learning algorithms now researcher make heart disease prediction models that are cheaper and more flexible instead of previous techniques [11], [12], [13]. in this prospectus several machine learning approaches using heart disease data sets for training and testing the particular model [14], [15], ,decision tree (al-qazzaz, mohammed et al. 2023), support vector machine are introduced by the researchers (javeed, rizvi et al. 2020)to predict heart disease . now by using these models this become very easy to predict heart disease timely. recent technology off deep learning enhances this field and increase the accuracy [15]. ii. methods researchers present eight different type of machine learning models to predict the heart disease. the models that they use for classification are decision tree, svm, xg boost algorithm, multinomial naive bayes algorithm, extra tree algorithm, logistic regression, adaboost and linear discriminant analysis algorithm. there method involves five steps, in step one they get data sets form different online data set collections, and secondly, they process the data by refining and standardization, the third step was hyper parameter tuning to get height accuracy by achieving hyper parameters best value. in step four they apply ml algorithm to classify. in this study the results shows that the accuracy is increased by using standardization of data set, in this study the accuracy of different classifier is improved up to 8.7 percent by standardization the data sets. the overall results support vector machine classifier results was best form all other classifier and it attain accuracy of 96.72% [16]. the authors use different ml algorithm to predict heart disease the algorithms include support vector machine, lr, gbc and knn. the use grid search cv with these algorithms. they use datasets form different source include beach v uci kaggle etc. the results of the study shows that the extreme gradient boosting algorithm along with grid search cv provide a best accuracy for both testing and training. the testing and training results was very good 100% and 99%. the study also highlights that using of optimal hyper parameter can enhance the performance of the algorithms [17]. this paper present different type of machine learning algorithms with feature selection algorithm to predict the heart disease. the algorithms that are used are knn, support vector machine, ld, gbc rf and dt, the algorithm that are used for feature selection is sequential feature selection. the results of the study shows that the random forest and decision tree provide more accurate results then others the results were respectively 100% and 99%. the study also illustrates that by using feature selection technique the accuracy of algorithms can be enhance [18]. in this research paper the researcher combines two different types of algorithms to propose a model for prediction of heart disease, the algorithms that they combine is random-forest & svm. the main objective to combine these algorithms was to eliminate the iterative feature for selecting the features for the disease. this was done to increase the accuracy of svm algorithm for diseases prediction. the vol. 6, no.1, january 2025 | 48 results of this research show that the hybrid model that they purpose gave more accuracy than the accuracy of individual algorithm support vector machine and random forest [19]. in this research the researcher uses different machine learning techniques to predict heart disease, they use k-nearest neighbor, rf, svm and multi-layer perceptron to attain their objective of predicting heart disease. the use different type of techniques such as evaluators of attribute, feature elimination and outlier removal to increase the overall performance of these algorithms. they use data sets for this purpose from different sources like cleveland, long beach v and uci kaggle for their model. the result of the heart disease prediction of this model was 82.47 to 100 percent [20]. the author presents a new hybrid approach to predict heart disease timely. they combine random forest and support vector machine for this purpose. they apply different techniques to risk factors. the researcher developer there model by using jupyter notebook online.the overall accuracy of this model show that the random forest classifier gave more accurate result then other algorithms [21]. this research article purpose sca_knn (sine cosine weighted k-nearest neighbour) ml algorithm to detect the heart disease in the patients. the model learns from the data that are stored in the block-chain. they use block-chain to ensure the integrity of the data of the patients. the results of the study shows that this purposed model gave more accuracy than the w k-nn & knn. the researcher also compares this algorithm with other algorithm in many ways like precision, mean square error. f1 score etc. this purposed algorithm and the storage system that they describe in their research have a good effect on increasing the accuracy of heart disease [22]. this study, the authors investigated five distinct methods (mmc, random, adaptive, quire, and audi) for deciding which data to include in a multi-label active learning scenario. by selecting the most crucial data points to query for their labels, the goal was to lower the expenses of labeling. to create predictive models for a dataset of heart disease, these selection techniques were paired with a label ranking classifier, and the classifier's hyperparameters were also tuned by using a grid search. overall, the study's findings point to the effectiveness of the selection approach in enhancing the learning model's accuracy beyond the data at hand when combined with the label ranking model. however, when comparing the models using the f-score, the performance of the selection process was very impressive when utilizing the optimum settings [23]. table 1. literature review reference model name overall accuracy remarks (absar, das et al. 2022) [24] random forest, adaboost, knn, decision tree rf: 99%, dt: 96%, ab: 100%, knn: 100% rf and dt have high accuracy, ab and knn achieve perfect accuracy, utilized streamlit for computeraided prediction system (yilmaz and yağin 2022) [25] random forest, support vector machine, logistic regression rf: higher accuracy utilized 10-fold repeated cross validation, rf model had higher accuracy and sensitivity (qu, deng et al. 2022) explainable boosting machine (ebm) 76% auc birth cohort investigation to predict congenital heart defects (chds), utilized ultrasound screening and ebm model (özbilgin, kurnaz et al. 2023) [26] support vector machine (svm) 93% accuracy utilized iris analysis and image processing for non-invasive cad diagnosis, demonstrated potential for early cad detection vol. 6, no.1, january 2025 | 49 (bhatt, patel et al. 2023) [27] multilayer perceptron, k-node clustering mlp: 87% utilized gridsearchcv for model optimization, achieved high accuracy with mlp, introduced knode clustering for improved accuracy (nandy, adhikari et al. 2023) [28] swarm-ann strategy 95.78% proposed swarm-ann strategy for smart healthcare framework, achieved high accuracy and outperformed conventional methods (manimurugan, almutairi et al. 2022) [29] hybrid linear discriminant analysis, faster rcnn hlda-malo: 96.85%, seresnext-101: 98.06% achieved high accuracy for sensor and echocardiogram classification, outperformed other models in accuracy (malnajjar and abu-naser 2022) [30] deep learning model 100% developed model to identify heart disease symptoms, utilized melfrequency cepstrum coefficient (mfcc) for feature extraction a. datasets our data originates from a 1988 merger of four distinct datasets: the long beach v dataset, the cleveland dataset, the hungarian dataset, and the swiss dataset. this dataset comprises various characteristics, with the "target" property indicating cardiac disease presence—0 for no illness and 1 for disease. the dataset includes vital information such as age, gender, chest pain type (categorized into four values), resting blood pressure, serum cholesterol level (mg/dl), fasting blood sugar status (binary for levels above 120 mg/dl), and resting electrocardiographic results (0, 1, 2). additional attributes include the maximum heart rate achieved during observation, signs of exercise-induced angina, the slope of the peak exercise st segment, the number of major vessels colored by fluoroscopy (0 to 3), and thalassemia status (normal, fixed defect, reversible defect). figure 1. correlational heatmap for heart stroke dataset vol. 6, no.1, january 2025 | 50 b. preparing data we first eliminate the missing values to prepare quality and reliable data. we removed six items in the cleveland dataset due to missing data. so the record reduces from 331 to 297 records. in the following iterations, we focused on shrinking the multiclass values of the predicted attribute, the presence of heart disease, into binary values. for this transformation, we took a value of 0 to indicate the absence of hd and 1 to indicate the presence of hd. we then standardized the data by converting all diagnostic values from 2 to 4 into 1, thus increasing the number of our dataset. in this way, the dataset had the quantity of diagnostic values set to 0 or 1, where 0 meant a lack of hd and 1 meant existence. the two main parts of the data set are characteristics and the target variable for developing our predictive model for heart diseases. standard independent variables that can be featured include age, gender, type of chest pain experienced, and other physiology-related parameters. these attributes are given as input to our model so that it can learn and give predictions. but the dependent variable, or target variable, will be the output for which we are trying to make the prediction: whether a patient has heart disease. decoupling features from the target variables is another essential part of preparing the dataset for training and evaluation with models. this would otherwise be an oversight. it enables us to feed the correct data into our model so that we can test the predictions with accuracy. after normalization and pre-processing, the data set must be divided into training and testing sets. the 80:20 popular approach splits the heart disease data set. eighty percent of the data is used to train the model, and twenty percent to test the model. by using the training set, data is given to our model for learning from it and the testing data set is applied to test the model for good performance. c. multilayer perceptron (mlp) multilayer perceptron is a primary type of artificial neural network, including layers of interacting nodes or neurons. it consists of one input layer, one or various hidden layers, and one output layer. there are built-in capabilities to learn complex patterns in the data and relationships among them. in an mlp, each neuron receives input, performs a linear combination, applies a non-linear activation function, and then passes the results to the next layer. mlps have been used in nearly all fields of classification and regression as well as pattern recognition problems since they can theoretically approximate any function that contains discontinuities with appropriate data and computational resourcing. in the context of our heart disease prediction, mlps can learn from these features to classify patterns in the data as a sign of the presence or absence of heart disease. d. evaluation measures various emulation measures were used to determine how much the model could predict heart diseases. these are accuracy, precision, recall, and f1-score. accuracy: percentage of cases that were classified correctly relative to the total number of cases. a model is considered dependable in predicting cardiac problems, and the more accurate it is. 𝐴𝑐𝑐 = 𝑇𝑃 + 𝑇𝑁 𝑇𝑃 + 𝑇𝑁 + 𝐹𝑃 + 𝐹𝑁 (1) precision: the accuracy of the model is measured in terms of how many positive instances have been correctly predicted by the total amount of positive cases. the exact identification of people with heart disease helps avoid many false positives. 𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛 = 𝑇𝑃 𝑇𝑃 + 𝐹𝑃 (2) recall: sensitivity, also called recall, measures the ratio of true positives that are successfully anticipated to actual positives. it shows the model's consistency in the wrong detection of heart problems. vol. 6, no.1, january 2025 | 51 𝑅𝑒𝑐𝑎𝑙𝑙 = 𝑇𝑃 𝑇𝑃 + 𝐹𝑁 (3) f1-score: it is a metric of evaluation for a model, although its computation involves recall and accuracy. given that both kinds of wrong outcomes, positive and negative, are taken into account, this would be a good test case for models on imbalanced datasets. 𝐹1 − 𝑆𝑐𝑜𝑟𝑒 = 2 × 𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛 × 𝑅𝑒𝑐𝑎𝑙𝑙 𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛 + 𝑅𝑒𝑐𝑎𝑙𝑙 (4) iii. results and discussions the way toward an adequate heart disease prediction has been marred with searching for such patterns and rigorous evaluation of diverse methodologies. in this work, we tried to take out the mystery of the role of machine-learning methods in the detection of subtle patterns that lead to the presence of heart diseases while focusing on the factors responsible for predictive accuracy and reliability. table 2. contribution results for thew heart stroke prediction with mlp 0 1 macro avrg weighted avg precision 0.97 1.0 0.99 0.99 recall 1.0 0.97 0.99 0.99 f1 score 0.99 0.99 0.99 0.99 accuracy 0.99 0.99 0.99 0.99 a. discussion we point out key insights gained through our investigation and what these mean for future research and clinical practice. one of the best performers in classification is a multilayer perceptron (mlp) model, which achieved an accuracy rate of 99%. the suggested model is strong and dependable, as shown by the assessment metrics for the models used to diagnose heart disease, which consistently show great performance across all measurements. the model was able to get accuracy values between 0.97 and 1.0 across several test sets, to begin with. nearly all occurrences that were categorised as positive were true, according to precision, which assesses the accuracy of the model's positive predictions. important in medical diagnostics for avoiding needless treatments or further invasive procedures, this high accuracy implies that the model is very good at reducing false positives. figure 2. a) training loss and validation loss is plotted over graph b) training accuracy is plotted against validation accuracy vol. 6, no.1, january 2025 | 52 the model's recall values were 1.0 and 0.97. according to recall, which is a measure of the model's accuracy in identifying real positive occurrences, the model does a great job of catching almost all genuine cases of heart disease. to guarantee that no occurrence of the illness goes unnoticed, and that treatment can begin promptly, high recall is especially critical in medical settings. the f1 score, which is a harmonic mean of recall and accuracy, remained constant at 0.99. the model's ability to accurately detect instances of heart disease while also minimising false positives is confirmed by this score, which strikes a balance between recall and precision. the model's balanced performance and its efficacy in sustaining accuracy under varied situations are further shown by the constancy of the f1 score across several test sets. the model's accuracy was 0.99 across the board, which means that almost all the predictions, positive and negative, were spot on. a high level of accuracy verifies that the model is suitable for practical use in clinical settings by demonstrating its overall dependability in producing accurate predictions. the model's generalizability and robustness are shown by its equal correctness across assessments. these features are necessary for a diagnostic tool that is meant to be used in numerous real-world circumstances. the findings show that the suggested model for diagnosing heart disease is quite effective in terms of accuracy, precision, recall, and f1 score. this impressive performance indicates that the model is not only good at detecting instances of heart disease, but also trustworthy in preventing false positives, guaranteeing thorough and precise diagnoses. since this is the case, the model may be relied upon by medical practitioners to reliably diagnose cardiac illness. that very high sensitivity level shows the strength of mlp in correctly defining cases of heart disease. furthermore, mlp has represented good accuracy, precision, recall, and f1-score metrics with the precision of classifying an individual with or without heart disease, which would be remarkable. figure 3. confusion matrix for all the measures, these models showed strong performance, which strengthens the effectiveness of collaborative intelligence in raising predictive accuracy and reliability. b. comparative analysis recent developments in machine learning and deep learning are shown by the significant discrepancies in accuracy found when comparing different models for the detection of heart disease. in vol. 6, no.1, january 2025 | 53 their study [38], found that traditional models like logistic regression (lr), k-nearest neighbours (knn), support vector machine (svm), random forest (rf), decision tree (dt), and a general deep learning (dl) approach had accuracies ranging from 83.3% to 94.2%. the deep learning model is the most impressive of the bunch, with an accuracy rate of 94.2%. on the other hand, ensemble approaches show that they are more effective. as to the findings of atallah and al-mousa [31], the hard voting ensemble model—which integrates many classifiers—attained a precision of 90.00%. better forecasts are produced by this strategy since it takes advantage of combining the capabilities of many models. while the ensemble methods outperform the individual classical models, the naive bayes (nb) classifier (84.51 percent accuracy) and the knn classifier (85 percent accuracy) also perform comparably, according to [32] and [33], respectively. the accuracy of 88.70% achieved by [34] when decision tree and random forest models were combined shows the effectiveness of ensemble approaches in improving prediction accuracy. additionally, anns have been investigated; however, [35] reported an accuracy of 82.49% using anns, suggesting that neural network topologies for the detection of cardiac disease need additional optimisation [36]. demonstrated an accuracy of 88.70% using a linear model and random forest, demonstrating the efficacy of hybrid techniques. according to what [37] stated, one remarkable model, lofs-ann (local outlier factor-support artificial neural network), managed to reach an accuracy level of 90.5%. this methodology improves neural networks' forecasting abilities by using anomaly detection. the suggested model in this research achieves a remarkable 99% accuracy, far surpassing all the preceding models. significant advancements in model design and training approaches, maybe using state-of-the-art techniques like convmixer for effective feature extraction and classification, are indicated by this. all things considered, the comparison study shows how heart disease diagnostic models have progressed, and the suggested model is the most accurate one yet for this vital medical application. table 3. comparison of contributing results with previous studies model accuracy reference lr, knn, svm, rf, dt, dl 83.3%, 84.8%, 83.2%, 80.3%, 82.3%, 94.2% (bharti, khamparia et al. 2021) hard voting ensemble 90.00% (atallah and al-mousa 2019) [31] nb 84.51% (tougui, jilbab et al. 2020) [32] knn 85.00% (pawlovsky 2018) [33] rf+dt 88.70% (kavitha, gnaneswar et al. 2021) [34] ann 82.49% (almazroi, aldhahri et al. 2023) [35] rf with a linear model 88.70% (mohan, thirumalai et al. 2019) [36] lofs-ann 90.5% (goyal 2022) [37] proposed models 99% purposed model iv. conclusion in summary, our research represents one giant stride toward realizing the transformational potential of machine learning in predicting heart diseases. we have demonstrated the efficacy of machine learning models for augmenting traditional approaches in heart disease diagnosis and risk assessment through careful experimentation, rigorous evaluation, and nuanced interpretation. we unravelled some novel insights into the complex interplay of factors contributing to heart disease manifestation and progression by applying state-of-the-art methodologies and a multidisciplinary approach. our results are, therefore, more than a technical tour de force; they also demonstrate the transformational impact machine learning in healthcare can enable for proactive and personalized healthcare interventions. as such, the promise of further increasing predictive accuracy, reliability, and interpretability with vol. 6, no.1, january 2025 | 54 continued exploration and innovation in machine learning techniques lies in these critical dimensions: ultimately advancing improved patient outcomes and better clinical decision-making. references [1] heidenreich, p. a., et al. (2011). "forecasting the future of cardiovascular disease in the united states: a policy statement from the american heart association." circulation 123(8): 933-944. [2] conrad, n., et al. (2018). "temporal trends and patterns in heart failure incidence: a populationbased study of 4 million individuals." the lancet 391(10120): 572-580. [3] lópez-sendón, j. (2011). "the heart failure epidemic." medicographia 33(4): 363-369. [4] allen, l. a., et al. (2012). "decision making in advanced heart failure: a scientific statement from the american heart association." circulation 125(15): 1928-1952. [5] ghwanmeh, s., et al. (2013). innovative artificial neural networks-based decision support system for heart diseases diagnosis. [6] al-shayea, q. k. (2011). "artificial neural networks in medical diagnosis." international journal of computer science issues 8(2): 150-154 [7] gavhane, a., et al. (2018). prediction of heart disease using machine learning. 2018 second international conference on electronics, communication and aerospace technology (iceca), ieee. [8] chen, k., et al. (2018). heart murmurs clustering using machine learning. 2018 14th ieee international conference on signal processing (icsp), ieee. [9] islam, a. m. and a. majumder (2013). "coronary artery disease in bangladesh: a review." indian heart journal 65(4): 424-435. [10] melnyk, j., et al. (2015). "awareness and knowledge of cardiovascular risk through blood pressure and cholesterol testing in college freshmen." american journal of health education 46(3): 138-143. [11] nahar, j., et al. (2013). "computational intelligence for heart disease diagnosis: a medical knowledge driven approach." expert systems with applications 40(1): 96-104. [12] sree, s. v., et al. (2012). "cardiac arrhythmia diagnosis by hrv signal processing using principal component analysis." journal of mechanics in medicine and biology 12(05): 1240032. [13] liu, n., et al. (2012). "an intelligent scoring system and its application to cardiac arrest prediction." ieee transactions on information technology in biomedicine 16(6): 1324-1331. [14] exarchos, k. p., et al. (2014). "a multiscale approach for modeling atherosclerosis progression." ieee journal of biomedical and health informatics 19(2): 709-719. [15] wiharto, w., et al. (2016). "intelligence system for diagnosis level of coronary heart disease with k-star algorithm." healthcare informatics research 22(1): 30. [16] saboor, a., et al. (2022). "a method for improving prediction of human heart disease using machine learning algorithms." mobile information systems 2022(1): 1410169. [17] ahmad, g. n., et al. (2022). "efficient medical diagnosis of human heart diseases using machine learning techniques with and without gridsearchcv." ieee access 10: 80151-80173. [18] ahmad, g. n., et al. (2022). "comparative study of optimum medical diagnosis of human heart disease using machine learning technique with and without sequential feature selection." ieee access 10: 23808-23828. [19] suresh, t., et al. (2022). "a hybrid approach to medical decision-making: diagnosis of heart disease with machine-learning model." international journal of electrical and computer engineering (ijece) 12(2): 1831-1838. [20] pal, m., et al. (2022). "risk prediction of cardiovascular disease using machine learning classifiers." open medicine 17(1): 1100-1113. vol. 6, no.1, january 2025 | 55 [21] ansarullah, s. i., et al. (2022). "significance of visible non‐invasive risk attributes for the initial prediction of heart disease using different machine learning techniques." computational intelligence and neuroscience 2022(1): 9580896. [22] hasanova, h., et al. (2022). "a novel blockchain-enabled heart disease prediction mechanism using machine learning." computers and electrical engineering 101: 108086. [23] el-hasnony, i. m., et al. (2022). "multi-label active learning-based machine learning model for heart disease prediction." sensors 22(3): 1184. [24] absar, n., et al. (2022). the efficacy of machine-learning-supported smart system for heart disease prediction. healthcare, mdpi. [25] yilmaz, r. and f. h. yağin (2022). "early detection of coronary heart disease based on machine learning methods." medical records 4(1): 1-6. [26] özbilgin, f., et al. (2023). "prediction of coronary artery disease using machine learning techniques with iris analysis." diagnostics 13(6): 1081. [27] bhatt, c. m., et al. (2023). "effective heart disease prediction using machine learning techniques." algorithms 16(2): 88. [28] nandy, s., et al. (2023). "an intelligent heart disease prediction system based on swarm-artificial neural network." neural computing and applications 35(20): 14723-14737. [29] manimurugan, s., et al. (2022). "two-stage classification model for the prediction of heart disease using iomt and artificial intelligence." sensors 22(2): 476. [30] malnajjar, m. k. and s. s. abu-naser (2022). "heart sounds analysis and classification for cardiovascular diseases diagnosis using deep learning." [31] atallah, r. and a. al-mousa (2019). heart disease detection using machine learning majority voting ensemble method. 2019 2nd international conference on new trends in computing sciences (ictcs), ieee. [32] tougui, i., et al. (2020). "heart disease classification using data mining tools and machine learning techniques." health and technology 10(5): 1137-1144. [33] pawlovsky, a. p. (2018). an ensemble based on distances for a knn method for heart disease diagnosis. 2018 international conference on electronics, information, and communication (iceic), ieee. [34] kavitha, m., et al. (2021). heart disease prediction using hybrid machine learning model. 2021 6th international conference on inventive computation technologies (icict), ieee. [35] almazroi, a. a., et al. (2023). "a clinical decision support system for heart disease prediction using deep learning." ieee access. [26] mohan, s., et al. (2019). "effective heart disease prediction using hybrid machine learning techniques." ieee access 7: 81542-81554. [37] goyal, s. (2022). predicting the heart disease using machine learning techniques. ict analysis and applications: proceedings of ict4sd 2022, springer: 191-199. [38] bharti, r., et al. (2021). "prediction of heart disease using a combination of machine learning and deep learning." computational intelligence and neuroscience 2021(1): 8387680. vol. 5, no.2 june 2024 | 64 identification of socio economic registration data using ocr based tesseract and google cloud vision lionardi ursaputra pratama 1*, aviv yuniar rahman 2, rangga pahlevi putra 3 1 2 3 department of informatics engineering, faculty of engineering, universitas widyagama malang, indonesia e-mail: ursaputra@cordelia.id*, aviv@widyagama.ac.id, rangga@widyagama.ac.id abstract: the indonesian government program, called socio-economic registration (regsosek), aims to measure and monitor the socio-economic conditions of low-income people. one of the relevant data used for research is regsosek. this method is used to analyze the influence of economic and social infrastructure on economic growth, analyze the socio-economic determinants of ownership of work accident insurance for informal workers, create a women's socio-economic vulnerability index (iksep), and study intercultural literacy from a social, economic and political perspective. the success of the government's socio-economic registration program depends on the role of data collection officers or surveyors, who directly interact with the community to obtain information about socio-economic registration (regsosek) data collection. this method also has other obstacles that significantly affect the overall results of the survey, where the survey results must be entered manually by the surveyor from a form with handwritten data, after which it is entered into the website. this method is vulnerable to human error, where the handwriting is difficult to read, and mistakes are made during the data input. the technology that can be used to handle this problem is implementing the ocr method, where writing that was initially handwritten manually can be identified and converted into digital text that can be edited (editable text) and processed automatically. this research shows that the proposed method has good accuracy, with an accuracy of 96.45%, cer 0.3%, and wer 4.30%. keywords: hand writing, optical character recognition, socioeconomic infrastructure, surveyor officer. i. introduction the indonesian government program, called socio-economic registration (regsosek), aims to measure and monitor the socio-economic conditions of low-income people. according to this program, social protection is the government's way of dealing with poverty in indonesia [14]. one of the relevant data used for research is regsosek. this method is used to analyze the influence of economic and social infrastructure on economic growth, analyze the socio-economic determinants of ownership of work accident insurance for informal workers, create a women's socio-economic vulnerability index (iksep), and study intercultural literacy from a social, economic and political perspective [11]. the success of the government's socio-economic registration program depends on the role of data collection officers or surveyors, who directly interact with the community to obtain information about socio-economic registration (regsosek) data collection. so that the public as respondents can provide answers that are appropriate to the conditions, this must be explained well in the field. however, there are obstacles to getting transparent information from respondents. this method also has other obstacles that significantly affect the overall results of the survey, where the survey results must be entered manually by the surveyor from a form with handwritten data, after which it is entered into the website. this method is vulnerable to human error, where the handwriting is difficult to read, and mistakes are made during the data input [12]. the technology that can be used to handle this problem is implementing the ocr method, where writing that was initially handwritten manually can be identified and converted into digital text that can be edited and processed automatically (memon et al., 2020). before the system can recognize human p-issn: 2715-2448 | e-issn: 2715-7199 vol.5 no.2 june 2024 buana information technology and computer sciences (bit and cs) vol. 1, no.2 june 2024 | 65 handwritten image patterns, information representing the image must be retrieved, known as input data. a scanning process is carried out on the resulting image to obtain digital data, and then a preprocessing process is carried out [7]. this process can complete the creation of an intelligent computer system that can recognize handwriting. often, these methods are combined with other algorithms to create new applications that can solve more complex problems [5]. currently, ocr is applied in various fields, including data entry. previous research utilized ocr as automation for logistics warehouse data input management [2]. the results of the introduction of ocr are used as input for the logistics warehouse management system, which will then be processed by the management system and entered into the database. in detecting handwriting, this research produces an accuracy of 46.9%. it shows that tesseract can be implemented to convert images into text. however, the recognition results depend on the quality of the image and the variety of text used as test data and taking images requires sufficient light. other research uses tesseract ocr for text recognition in social survey data. this research produced a cer of 2.60% and a wer of 25.20% using the tesseract method [1]. the difference in accuracy and error rate is due to the library's dependence on the test document's quality, so the ocr output is often inappropriate. both studies show that the quality of the image or document used as a source scanned by an ocr application is the main problem in using ocr on handwritten text. tesseract is an open-source ocr application created by hp from 1984 to 1994. tesseract was first created as a graduate project and released by hewlett packard and the university of nevada, las vegas, in 2005. now, google partially funds its development, and version 2.0, tesseract 4. x, was released in december 2019. tesseract has four main functional modules: static character classifier, word recognition, paragraph and sentence search, linguistic analysis, and adaptive classifier. however, if the image is preprocessed, the results from tesseract will be much better. to overcome this problem, research was carried out to automate data input in handwriting from survey forms using the ocr method based on google vision and tesseract. this research aims to create a system that can be used to recognize handwriting so that it can be easily recognized as text using computer vision methods. ii. methods this research is divided into several general stages. this stage includes literature study, planning, data collection, implementation, testing, and report preparation. these stages will be carried out sequentially to produce an optimal report. the general stages of research can be seen in figure 1. the literature study phase involves a comprehensive review of existing research, theories, and relevant publications to establish a solid theoretical framework for the study. fig1. study method a. literature study in this research, the references used come from books, journals, ebooks, previous research and other reliable sources. the literature study used google scholar as the primary search tool for references related to text detection, tesseract and google vision. literature research was carried out to obtain information about several main topics of this research, namely: vol. 5, no.2 june 2024 | 66 a) ocr b) google cloud vision c) tesseract d) merged method b. data gathering the required data collection stage is the most important in this research. the data collected is used for system testing, whether the method used is optimal for the test data. the survey form image data was collected by scanning 100 data. c. ocr implementation based on figure 2, this research began by entering the survey data image. after that, the image is processed for thresholding and greyscale. this pre-processing is done so ocr can more easily differentiate between objects and backgrounds. next is detecting text; tesseract and google vision are used to extract images into asci characters. after the asci characters are obtained, compare the results of the two methods, and which is the best with ground truth? then, the best results will be normalized and entered into the ignore list process to determine which words will be used and which will be ignored. the following process is testing, which will use the cer, wer and accuracy parameters. these metrics will assess the performance of the ocr methods and provide valuable insights into their effectiveness in accurately converting the images into editable text, thereby ensuring the reliability of the data obtained from the survey images. calculations of accuracy, word error rate (wer), and coefficient of error rate (cer) can be made using formulas 1, 2, and 3, respectively. cer = (insertion + removal + substitution)/totalnumberofcharactersingroundtruth (1) 𝑊𝐸𝑅 = (𝐼𝑛𝑠𝑒𝑟𝑡𝑖𝑜𝑛 + 𝐷𝑒𝑙𝑒𝑡𝑖𝑜𝑛 + 𝑆𝑢𝑏𝑠𝑡𝑖𝑡𝑢𝑡𝑖𝑜𝑛)/𝑇𝑜𝑡𝑎𝑙𝑛𝑢𝑚𝑏𝑒𝑟𝑜𝑓𝑤𝑜𝑟𝑑𝑠𝑖𝑛𝐺𝑟𝑜𝑢𝑛𝑑𝑇𝑟𝑢𝑡ℎ (2) 𝐴𝑐𝑐𝑢𝑟𝑎𝑐𝑦 = (1– 𝑎𝑣𝑒𝑟𝑎𝑔𝑒 𝐶𝐸𝑅) ∗ 100 (3) fig 1. optical character recognition process these assessments were fundamental in ensuring the fidelity and precision of the data derived from the survey images, thereby bolstering the overall quality and trustworthiness of the research outcomes. iii. results a. literature study a. optical character recognition (ocr) optical character recognition (ocr) is an application that can recognise ascii characters in digital photographs and turn them into editable text data [6]. as seen in figure 3, the ocr procedure often comprises multiple steps. the three main processes in this procedure are preprocessing, feature extraction, and recognition. preprocessing entails improving vol. 1, no.2 june 2024 | 67 the image and reducing noise, feature extraction finds essential elements, and recognition converts these elements into text [9]. fig 2. ocr process flowchart b. tesseract tesseract is a widely used optical character recognition (ocr) library for processing scanned images for text recognition. this engine uses neural networks to analyze image patterns and recognize characters with high accuracy. this library can recognize characters in various languages, such as arabic, bulgarian, catalan, chinese, czech, danish, dutch, english, french, german, greek, hindi, indonesian, and italian. some studies have suggested using convolution-based preprocessing with specific kernels to improve the accuracy of the tesseract ocr engine [3]. fig 3. architecture of tesseract library first, the text area in the image is identified through page layout analysis. next, the found text areas are divided into several "blobs". blobs are classifiable units consisting of multiple characters or parts of multiple characters [10]. the third step is to determine the lines of text and combine the blobs into a series of words that fill the space. the first step is to prepare the words for recognition. the next step is to recognize each word in two routes. tesseract breaks down and merges the single word in each path into blobs, forming a series of identifiable character outlines. tesseract recognizes the character vol. 5, no.2 june 2024 | 68 outline on the first computer with a static classifier based on the feature library. the architecture of tesseract is depicted in figure 4. c. google vision fig 4. flowchart of google vision google vision is a computer vision technology developed by google that uses machine learning to analyze and understand visual information in images and videos. this technology offers many features, such as image processing, face processing, object detection, and labelling. the google cloud vision api allows developers to integrate this service into their applications [4]. the utilization of google vision can be implemented, as shown in figure 5. b. implementing method the end product of this research is an application that can be used to extract text from handwritten text photographs. python is the programming language used to create the program, and it runs on the google colab platform. google colab produces an excel file with the file name and detection results as part of its output. and the outcomes of detection. the excel file's detection findings are anticipated to make it easier for surveyors to interpret handwritten responses on survey sheets. the user will be prompted to input the image directory and the output file name in the main program, which is an excel file. figure 6 displays the program display. fig 5. input and output program setting up the google cloud vision api service credentials using the google_application_credentials environment variable completes the procedure. next, the input directory's image file list (img_files) is initialised. one method of obtaining truth data is using an excel file, previously mentioned in ground_truth_file. the file's text data is loaded into a pandas data frame, and the 'text' column is adjusted to change the text to lowercase (lowercase). then, in identifying text in photos, a list of words (ignore_list) that will be disregarded includes pointless or useless terms. every picture file in the img_files directory undergoes one iteration. tesseract ocr with google cloud vision api's ocr (optical character recognition) technique extracts text from each read image. additionally, word similarity between the resultant text and the truth text (ground_truth_text) was determined using the natural language toolkit (nltk) library's cer (character error rate) and wer (word error rate) metrics. the text detection findings are displayed as graphics with distinct colour markers for each found word. for every word in ground_truth_df, an iteration loop completes the process. the average is then determined by adding up the findings of the cer and wer measurements. the application then uses a pandas dataframe-which will be added or built if it does not exist-to store the detection results in an excel file. vol. 1, no.2 june 2024 | 69 the print programme averages the cer and wer from the tesseract technique, google vision, and a mixture of both after all photos have been analysed. the accuracy of each approach is then printed by the programme as well. this program records and saves the detection findings into an excel analysis file for additional study. it also gives a comprehensive report on the accuracy of the text detection results from the two ocr algorithms utilised. figure 7 displays certain outcomes from the programme output. fig 6. output results and pre-processing the ignore list condition is used in the subsequent findings, as shown in table 1. during the preprocessing step, the otsu approach gives the image better contrast, eliminates noise, and facilitates program processing. all the text in the image is readable correctly when viewed from the output the programme generated using the suggested way. table 1. text extraction result ground truth ignore list result 101. provinsi 102. kabupaten /kota *) 103. kecamatan 104. desa/kelurahan*) 105. kode sls/non sls 106. nama sls/non sls 107. alamat (jalan/gang, nomor rumah) jawa timur jawa timur 35 35 malang malang 73 73 blimbing blimbing 040 040 purwodadi purwodadi 008 008 0057 0057 kode kode sub sls sub sls 00 00 rt. 007 rt.007 rw. 008 rw.008. jl plaosan barat 83 ji plaosan barat 83 in comparing image processing methods as shown in table 2, data 1 uses a more straightforward image processing technique: grayscale on the tested images. meanwhile, in data 2, the method applied vol. 5, no.2 june 2024 | 70 is more sophisticated, applying the otsu thresholding method. otsu thresholding is a technique used to segment images by determining thresholds automatically based on image histogram analysis. using grayscale in data 1 means only using the brightness channel of the image, ignoring colour information. although simple, grayscale can be helpful in some cases, especially when colour accuracy is not the central aspect required in text or object recognition. table 2. otsu tresholding implemented and grayscale only file name cer difference wer difference accuracy difference data-1.jpg 0.23% (increase) 1.05% (increase) -0.23% (decrease) data-2.jpg -0.27% (decrease) -0.42% (decrease) 0.27% (increase) data-3.jpg 0.18% (increase) 0.21% (increase) -0.18% (decrease) data-4.jpg -0.08% (decrease) 0.20% (increase) 0.08% (increase) data-5.jpg 0.13% (increase) -0.42% (decrease) -0.13% (decrease) data-6.jpg -0.35% (decrease) -0.85% (decrease) 0.35% (increase) data-7.jpg 0% 0% 0% data-8.jpg 0% 0.42% (increase) 0% data-9.jpg 0.23% (increase) -0.19% (decrease) -0.23% (decrease) data-10.jpg 0.29% (increase) 0.42% (increase) -0.29% (decrease) on the other hand, applying otsu thresholding to data 2 can provide advantages because this method can automatically determine the optimal threshold for separating objects from the background. the table shows the results of a comparative analysis of the performance of tesseract and google vision in terms of character error rate (cer), word error rate (wer), and accuracy differences between ten different data pictures for optical character recognition (ocr). the two approaches in each photograph show apparent differences in performance metrics. for example, tesseract and google vision both show higher cer and wer in data-1.jpg, which is a factor in tesseract's decreasing accuracy. on the other hand, google vision keeps its cer constant while somewhat increasing its wer. the combined data show a modest loss in overall accuracy along with a minor rise in cer and wer. both approaches show a drop in cer and wer in data-2.jpg, with tesseract showing a more significant improvement over google vision. as a result, tesseract's accuracy increases noticeably, while google vision's accuracy slightly increases. the combined results show a drop in wer and cer; however, accuracy has somewhat decreased. from data-3.jpg to data-10.jpg, the remaining photos show consistent patterns of volatility. specifically, differences in cer, wer, and accuracy point to distinct advantages and disadvantages between tesseract and google vision on various datasets. using this technique, images may undergo better segmentation, improving the system's ability to better recognize text or objects in the image. otsu method the otsu method uses discriminant analysis to find variables that distinguish between two or more naturally occurring groups [15]. discriminant analysis maximizes these variables to divide objects into foreground and background. the discriminant analysis produces a threshold value that divides a grayscale image into two groups of black and white [13]. c. analysis the combination approach proved to be more effective than either one alone. while tesseract and google vision have somewhat higher average cers-roughly 1.8% to 1.9% for tesseract and 1.0% to 1.7% for google vision-the combined method's average cer is 1.8% to 1.9%. for the same reason, the combined approach performs better when calculating the word error rate (wer). the combined method's average wer ranges from 1.9% to 2.0%, tesseract's from 2.0% to 2.5%, and google vision's from 1.7% to 2.2%. the error rate in text detection in photos is successfully decreased by combining the output of two ocr algorithms. this demonstrates that using google vision in conjunction with tesseract ocr can yield better accurate results regarding text recognition on photos than either technique alone. vol. 1, no.2 june 2024 | 71 as a result, the program's output results demonstrate that the combined approach can identify text on images with a greater accuracy rate, making it a better option for text recognition applications requiring a low error rate. discussion reference novelty results of previous research results of the proposed method (arianto et al., 2023) difference: previous research produced a cer of 2.60% and a wer of 25.20% in detecting text and did not include the accuracy obtained in the test. we investigated how to reduce the cer and wer values and obtain high accuracy in detecting written text. hand. novelty: this research shows that the combination of ocr using google vision and tesseract is a method with a lower wer level, and there is a decrease in cer and wer compared to before. cer: 2.60% wer: 25.20% accuracy: cer: 0.3% wer: 4.30% accuracy: 96.45% (berg, so and seo, 2019b) difference: this research only produces an accuracy of 46.9% in detecting handwriting, we analyze how accuracy in detecting handwriting can be improved. novelty: this study shows that the most accurate ocr approach and also shows improvements compared to previous methods is the combination of tesseract and google vision. cer: wer: accuracy: 46.9% (thammarak et al., 2022) difference: this research only produces accuracy on tesseract of 47.02% and produces an accuracy of 84.03% in experiments using google vision in detecting writing. we analyze how accuracy in detecting handwriting tesseract cer: wer: accuracy: 47.02% google vision cer: wer: accuracy: 84.03% vol. 5, no.2 june 2024 | 72 can be increased by combining the two methods based on evaluations from previous research. novelty: this study shows that the most accurate ocr approach and also shows improvements compared to previous methods is the combination of tesseract and google vision. table 4.4 illustrates a comparison of the results of this study with previous research. previous research, without considering accuracy, recorded cer and wer for text detection of 2.60% and 25.20%. in an effort to achieve high levels of accuracy, as well as lower cer and wer, this study shows that utilizing tesseract and google vision together for ocr can reduce wer values and increase accuracy up to 96.45%. on the other hand, previous research only achieved an accuracy of 46.9% in identifying handwriting. however, the results of this research show that combining tesseract with google vision is the most accurate ocr solution and provides significant progress compared to previous methods. it is important to note that the collaboration of the two ocr methods not only results in a higher level of accuracy, but also provides significant progress compared to previous studies. these final results confirm that the combined approach between tesseract and google vision is not only effective in reducing wer values, but also creates a superior ocr solution in recognizing handwriting, overcoming the obstacles faced by previous methods. iv. conclusion the program uses tesseract ocr, google vision api, and combination approaches to recognise text on images. the study of the program's output results concludes that the combined method better recognises text on photos. while both google vision api and tesseract ocr have reasonable accuracy rates, combining the two results in a far lower mistake rate regarding text recognition. compared to individual approaches, the combined method successfully demonstrated lower average cer (character error rate) and wer (word error rate). this demonstrates how text identification accuracy in photos can be increased using a method that combines the outcomes of tesseract ocr with google vision. therefore, using integrated approaches can be a more efficient and best option for applications that need text recognition with a low mistake rate. this conclusion demonstrates that combining different techniques can yield more dependable and accurate results when it comes to text recognition in photos. references [1]. arianto, r. f., rahman, a. y., & marisa, f. (2023). text recognition for socioeconomic data survey sheet using ocr tesseract. [2]. berg, s. a., so, r. h. y., & seo, s. y. (2019). application of optical character recognition with tesseract in logistics management. international journal of internet manufacturing and services, 6(3), article 3. https://doi.org/10.1504/ijims.2019.10022461 [3]. chesley, e., marcantonio, j., & pearson, a. (2019). towards syriac digital corpora: evaluation of tesseract 4.0 for syriac ocr. hugoye: journal of syriac studies, 22(1), article 1. https://doi.org/10.31826/hug-2019-220105 [4]. gonzález, g., & evans, c. l. (2019). biomedical image processing with containers and deep learning: an automated analysis pipeline: data architecture, artificial intelligence, automated processing, containerization, and clusters orchestration ease the transition from data vol. 1, no.2 june 2024 | 73 acquisition to insights in medium‐to‐large datasets. bioessays, 41(6), article 6. https://doi.org/10.1002/bies.201900004 [5]. haji, c. m. (2022). linguistic analysis on cursive characters. the journal of duhok university, 25(2), article 2. https://doi.org/10.26682/sjuod.2022.25.2.3 [6]. huang, j., pang, g., kovvuri, r., toh, m., liang, k. j., krishnan, p., yin, x., & hassner, t. (2021). a multiplexed network for end-to-end, multilingual ocr. 2021 ieee/cvf conference on computer vision and pattern recognition (cvpr), 4545–4555. https://doi.org/10.1109/cvpr46437.2021.00452 [7]. hukkeri, g. s., goudar, r. h., janagond, p., & patil, p. s. (2022). machine learning in ocr technology: performance analysis of different ocr methods for slide-to-text conversion in lecture videos. international journal of advanced computer science and applications, 13(8), article 8. https://doi.org/10.14569/ijacsa.2022.0130839 [8]. memon, j., sami, m., khan, r. a., & uddin, m. (2020). handwritten optical character recognition (ocr): a comprehensive systematic literature review (slr). ieee access, 8, 142642– 142668. https://doi.org/10.1109/access.2020.3012542 [9]. muharom, s. (2019). pengenalan nomor ruangan menggunakan kamera berbasis ocr dan template matching. inform : jurnal ilmiah bidang teknologi informasi dan komunikasi, 4(1), article 1. https://doi.org/10.25139/inform.v4i1.1371 [10]. mursari, l. r., & wibowo, a. (2021). the effectiveness of image preprocessing on digital handwritten scripts recognition with the implementation of ocr tesseract. computer engineering and applications journal, 10(3), article 3. https://doi.org/10.18495/comengapp.v10i3.386 [11]. putri, m. h., & yuhan, r. j. (2020). indeks kerawanan sosial ekonomi perempuan indonesia tahun 2017. seminar nasional official statistics, 2019(1), article 1. https://doi.org/10.34123/semnasoffstat.v2019i1.117 [12]. rohman, m. a. a., & djasuli, m. (2022). penerapan good corporate governance tranparansi terhadap kinerja surveyor registrasi sosial ekonomi dalam mewujudkan data akurat. [13]. smith, r., newton, c., & cheatle, p. (n.d.). adaptive thresholding for ocr: a significant test. [14]. suharto, e. (2015). peran perlindungan sosial dalam mengatasi kemiskinan di indonesia: studi kasus program keluarga harapan. sosiohumaniora, 17(1), article 1. https://doi.org/10.24198/sosiohumaniora.v17i1.5668 [15]. wibawa, c., & anggraeni, d. t. (2023). comparison of image segmentation method in image character extraction preprocessing using optical character recoginiton. jurnal teknik informatika (jutif), 4(3), 583–589. https://doi.org/10.52436/1.jutif.2023.4.3.956 implementation of orange data mining to predict student graduation on time at pringsewu muhammadiyah university roby novianto1*, bambang triraharjo2, baskoro3 1,2,3 sistem dan teknologi informasi, universitas muhammadiyah pringsewu 1robynovianto@umpri.ac.id, 2bambangtriraharjo@umpri.ac.id, 3baskoro@umpri.ac.id abstract thel prolcelss olf molnitolring and elvaluating thel graduatioln olf muhammadiyah pringselwu univelrsity (umpri) studelnts relally ne lelds tol bel dolne l belcausel thel studelnt graduatioln ratel is an ellelmelnt olf accrelditatioln asselssmelnt that is velry impolrtant folr e lach study prolgram. data mining can be l use ld tol classify studelnt graduatioln accuracy. this study aims to l apply thel olrangel data mining applicatioln using thel k-nelarelst nelighbolr (k-nn), de lcisioln trelel and naivel bayels moldells and will theln elvaluatel thel accuracy olf elach olf thelsel moldells. this relselarch was colnducteld at pringselwu muhammadiyah univelrsity in selvelral batche ls, theln studelnt data will bel analyzeld using thel olrangel data mining applicatioln using thel k-nn, de lcisioln trelel and naive l bayels moldells. thel data te lsting prolcelss appliels k-folld cro lss validatioln (k=5), whilel thel e lvaluatioln moldell useld is thel colnfusioln matrix and rolc. thel relsults olf thel colmparisoln olf thel threlel moldells arel as folllolws, k-nn has an accuracy ratel olf 75.7%, delcisioln trelel has an accuracy ratel olf 78.1%, and naivel baye ls has an accuracy ratel olf 77.8%. thelrelfolrel, folr classifying thel graduatioln ratel olf muhammadiyah univelrsity studelnts, pringselwu relcolmmelnds thel de lcisioln trelel moldell belcausel it has a belttelr lelvell olf accuracy than k-nn and naivel bayels. kelywolrds: graduatioln, preldictioln, data mining, c4.5, naïvel bayels i. introduction the development of the world of education in indonesia has had the impact of very tight competition. this was triggered by the increasingly advanced education in universities. one of the impacts of competition is producing quality graduates. the criteria for quality graduates include being able to complete mass learning on time. the timely learning period greatly influences the quality of higher education [16]. the ability of universities to produce graduates who are able to solve learning problems on time is a factor that influences higher education accreditation. this is in accordance with the national accreditation board for higher education regulations number 3 of 2019 concerning higher education accreditation instruments which states that one of the indicators for accreditation assessment is the percentage of graduates on time for each program from a tertiary institution. therefore, it is necessary to monitor the student's study period. the average study period for students studying at pringsewu muhammadiyah university is still over 4 years so it is necessary to try an evaluation using the student classification method using the orange application with three models, namely k-nearest neighbor (knn), decision tree and naive bayes. length of study is the period of time required for students to complete their education. the duration of student study has been regulated in the ministry of education and culture's decree regarding the undergraduate program (s1) education system which has a semester credit load that must be taken between 144 and 160 credits with a length of study on campus of between 8 and 10 semesters or the equivalent of between 4 and 5 years. information on students' grades for each semester and student graduation information can be processed to create data that is useful for analyzing the accuracy of students' study progress [1]. based on information obtained from muhammadiyah pringsewa university, in the 2015 to 2020 class year with an average number of graduates of 63%, data obtained that the average student study period was p-issn: 2715-2448 | e-issn: 2715-7199 vol.5 no.1 januari 2024 buana information technology and computer sciences (bit and cs) vol. 5, no.1 | 29 mailto:robynovianto@umpri.ac.id mailto:bambangtriraharjo@umpri.ac.id mailto:baskoro@umpri.ac.id still over 4 years [6]. therefore, it is necessary to try to analyze the factors that support the punctuality and delays in the student's study period. previous research related to predicting student graduation and classification mostly used k-nn, svm, neural network and naive bayes modes. [2],[14]. on the other hand, many previous studies have reviewed the results of comparative analyzes of several data mining models used to classify poor people [1], umkm income classification [3], nutritional classification [8], as well as classifications for the agricultural industry and public health [12]. this research aims to classify the timeliness of the study period of muhammadiyah pringsewu university students by applying three methods, namely knn, naive bayes and decision tree. next, a comparative analysis of the three models will be carried out by applying confusion matrix and roc analysis to ensure the level of accuracy of the three methods. this research contribution really helps the management of muhammadiyah university of pringsewu to develop strategies to minimize students who are not on time in completing their studies and contributes to determining the accuracy performance of several data mining methods, including knn, naive bayes and decision tree. ii. methods 1. research flow this research aims to carry out a comparative analysis of the knn, naive bayes and decision tree methods used to classify graduating students at muhammadiyah university of pringsewu. the application used for the simulation is orange data mining, an open source data mining application that has been proven to be able to help researchers analyze their data. the process stages in this research can be seen in figure 1. figure 1 research flow according to figure 1, the first step is problem identification, formulation and literature review. this is done first to develop research objectives and research contributions [10]. second is the process of collecting data, namely compiling training data and test data as a source of data classification. third is the process of designing the orange data mining widget for the student graduation classification process and method comparison. fourth is the process of classifying graduating students from muhammadiyah university of pringsewu using the knn, decision tree and naive bayes models. fifth is the process of evaluating the performance of classification methods and analyzing the comparison results of these methods. a. k-nearest neighbor (k-nn) k-nearest neighbor (k-nn) is a supervised method which means it requires training information to classify objects that are very close. the working principle of k-nn is to find the shortest distance between the information to be evaluated and k neighbors in the training data [11]. the dataset is grouped manually according to the type of student data pringsewu muhammadiyah university. the dataset used as a reference is 35 student data for which the classification model will be tested. next, the formula calculates the similarity of the dataset vector to each training dataset that has been classified. the k-nn theorem for calculating distance universally is as follows: 𝑑𝑖 = √∑(𝑥1𝑗 − 𝑛 𝑖=1 𝑝𝑗) 2 problem identification, problem formulation and literature review dataset orange data mining design classification process using knn, naive bayes and decision tree methods evaluation of classification performance and method comparison results vol. 5, no.1 | 30 information: di = sample distance xij = knowledge sample data pj = data input var ke-j n = number of samples the stages of the process of implementing the k-nn method are as follows: 1) determine the parameter k (number of closest neighbors). 2) calculates the square of the object's euclidean distance to the given training data. 3) sort result number 2 in ascending order (in order from high to low) 4) collecting y categories (nearest neighbor classification based on k value) 5) by using the nearest neighbor category with the majority, the object category can be predicted. b. decision tree (c4.5) decision tree based on c4.5 algorithm is a commonly used classification technique to extract relevant relationships in data. the c4.5 algorithm is a program that creates a decision tree based on a labeled input data set. the advantage is that the model can be easily interpreted and implemented with both continuous and discrete values. the c4.5 algorithm divides training data with the help of information acquisition [17]. attributes that have high frequencies are considered to separate data based on the information available in the dataset. when calculating the gain value, you need to know the entropy value, namely using the following formula: 𝐸𝑛𝑡𝑟𝑜𝑝𝑦 (𝑖) = 𝑓(𝑖, 𝑗). 2𝑓 [(𝑖, 𝑗)] 1) gain value using the formula: 𝑔𝑎𝑖𝑛 = − . 𝐼𝐸 (𝑖) 2) to calculate the gain ratio, you need to know a new term called split information with the formula: 𝑆𝑝𝑙𝑖𝑡𝐼𝑛𝑓𝑜𝑟𝑚𝑎𝑡𝑖𝑜𝑛 = -∑𝑐 𝑡=1 𝑆1 𝑆 𝑙𝑜𝑔2 𝑆1 𝑆 3) next, calculate the gain ratio 𝐺𝑎𝑖𝑛𝑟𝑎𝑡𝑖𝑜 (𝑆, 𝐴) = 4) repeat step 2 until all records have been split. the decision tree splitting process ends when: 1) all tuples in node record m are of the same class. 2) the attributes in the dataset are not further divided. 3) an empty branch has no records c. naïve bayes bayesian classification is a statistical classification that can be used to predict the probability of membership of a class discovered by the british scientist thomas bayes [4]. naive bayes is a classification algorithm that is quite simple and easy to implement so this algorithm is very effective when tested with the correct data set, especially if naive bayes is combined with function selection, so naive bayes can reduce redundancies in the data, besides that naive bayes shows good results when combined with clustering methods. naive bayes is proven to have high accuracy compared to support vector machines. vol. 5, no.1 | 31 vol. 5, no.1 | 31 then x is evidence, h is hypothesis, p(h|x) is probability that hypothesis h is true, evidence of true or hypothesis h or posterior so variable c explains the class, while variables f1...fn explain the character of the instructions in carrying out the classification. where this formula explains the probability that the sample enters a special character in class c (posterior), namely the probability that it comes out of class c (before entering the sample, many priors are made), multiplied by the probability of the sample character appearing in class (also called likelihood), divided by the probability of the character appearing global examples (also called evidence). the formula above can be made simply as follows 𝑃𝑜𝑠𝑡𝑒𝑟𝑖𝑜𝑟 = 𝑃𝑟𝑖𝑜𝑟 𝑥 𝑙𝑖𝑘𝑒𝑙𝑖ℎ𝑜𝑑 𝑒𝑣𝑖𝑑𝑒𝑛𝑐𝑒 for continuous data classification, the gaussian density formula is used: where: p: opportunity xi: atributke i xi: the value of the i attribute y: class sought yi: subclass y is sought μ: mean, explains the average of all attributes σ: standard deviation, explained variance across attributes. 2. evaluasi kinerja a. confusion matrix this method only uses a matrix table as in table 1, if the dataset only consists of two classes, one class is considered positive and the other negative [5]. evaluation with the confusion matrix produces accuracy, precision and recall values. table 1 confusion matrix correct clalssificaltion clalssified als + + true positives fallse negaltives fallse positives true negaltives true positive is the number of positive records that are classified as positive, false positive is the number of negative records that are classified as positive, false negative is the number of positive records that are classified as negative, true negative is the number of negative records that are classified as negative, 𝐴𝐶𝐶 = 𝑇𝑃+𝑇𝑁 𝑇𝑃+𝑇𝑁+𝐹𝑃+𝐹𝑁 𝑃 = 𝑇𝑃 𝑇𝑃+𝐹𝑃 𝑆𝑛 = 𝑇𝑃 𝑇𝑃 + 𝐹𝑁 𝑆𝑝 = 𝑇𝑁 𝑇𝑁 + 𝐹𝑃 vol. 5, no.1 | 32 b. kurva roc the roc curve is a graphical plot that illustrates the diagnostic ability of a binary classifier system as its discrimination threshold varies [2]. this method was originally developed for military radar receiver operators starting in 1941, giving rise to its name. the roc curve is created by plotting the true positive rate (tpr) against the false positive rate (fpr) at various threshold settings. true positive rate is also known as sensitivity, recall, or probability of detection. the false positive rate is also known as the false alarm probability and can be calculated as (1 specificity). the roc can also be thought of as a plot of the power as a function of the type i error of the decision rule (when performance is calculated only from a sample of the population, it can be thought of as an estimator of this quantity). auc accuracy performance can be classified into several groups, namely [7]: 1. 0.90 – 1.00 = excellent classification 2. 0.80 – 0.90 = good classification 3. 0.70 – 0.80 = fair classification 4. 0.60 – 0.70 = poor classification 5. 0.50 – 0.60 = failure classification iii. results and discussions 1. data mining process in analyzing the performance of several classification models in the orange tool, a comparison of several data mining methods was carried out to select the best method with high accuracy, in classifying the pringsewu muhammadiyah university graduation status dataset as shown in figure 3. figure 3. widget design for student graduation status classification model in figure 3, a widget is designed using a classification model in orange data mining software in the form of k-nn, decision tree and naive bayes which is input by a dataset that has been previously processed. then the dataset is processed into classification mode. 2. classification model testing process in the process of testing the classification model that has been created previously, a collection of test data is needed to determine the classification results as shown in figure 4. vol. 5, no.1 | 33 figure 4. widget design for student graduation status dataset classification model in figure 4 is a widget design that has been added to the classification testing process for the classification model. in the red box image is a set of trial data that is entered into the classification process to find out the results of the graduation classification of muhammadiyah university pringsewu students. 3. evaluation process of classification model comparison results the next process is to carry out a classification model comparison process using test and score which is needed to calculate the success rate between each classification model in orange data mining as shown in figure 5. figure 5. widget design for calculating the success of the classification model in figure 5 is a widget design that has been added to the process of calculating the success rate of the classification model using the test and score widget, which will then be evaluated for accuracy using confusion matrix and roc analysis. 4. simulation results of 3 classification models the simulation results of the classification model were carried out using a test data set with 1 attribute as the target, 10 numeric attributes, namely nim, study program, gender, marital status, employment status, 3rd sem gpa, 4th sem gpa, 5th sem gpa, total sks, origin students, information, so that test score results are obtained as shown in figure 6. vol. 5, no.1 | 34 figure 6. test and score widget results based on student data that has been tested, the calculation results of precision, recall and accuracy for each model are obtained as shown in figure 6. model classification resultsk-nn, decision tree as well asnaive bayes shows that the accuracy valuedecision tree the highest is 78%. based on figure 6 which also shows a comparison of 3 auc models, it is known that the highest auc value is the methodk-nn namely 0.795. auc is used to measure discriminatory performance by estimating the probability of output from randomly selected examples from a positive or negative population. the greater the auc, the better the classification results used. 5. evaluation results with confusion matrix confusion matrix is a performance measurement for machine learning classification problems where the output can be in the form of 2 or more classes. confusion matrix is a table with 4 different mixtures of predicted values and actual values. the evaluation results for each classification model can be seen in figure 7 for the k-nn model, while the confusion matrix results for the decision tree model can be seen in figure 8 and the confusion matrix values for the naive bayes model can be seen in figure 9. figure 7. confusion matrix value of the k-nn method figure 7 shows that the value of true positive (tp) is 541, true negative (tn) is 154, false positive (fp) is 83, and false negative (fn) is 140. so the accuracy, precision and recall values of the k method -nn is as follows: vol. 5, no.1 | 35 figure 8. confusion matrix value for the decision tree method figure 8 shows that the value of true positive (tp) is 535, true negative (tn) is 182, false positive (fp) is 89, and false negative (fn) is 112. so the accuracy, precision and recall values of the decision method tree is as follows: figure 9. confusion matrix value of the naive bayes method figure 9 shows that the value oftrue positif (tp) is 552, true negative (tn) is 162false positive (fp) is 72, andfalse negative (fn) is 132. then the valueaccuracy, precision dan recall from the methodnaive bayes are as follows: based on the results of evaluation and validation usingconfusion matrix comparative values are obtainedaccuracy, precision dan recall of 3 methodsk-nn, naive bayes, and decision tree as seen in table 2. vol. 5, no.1 | 36 table 2. performance comparison metode accuracy precision recall k-nn 75,7% 74,8% 75,7% nalive balyes 77,8% 77% 77,8% decision tree 78,1% 77,7% 78,1% based on table 2, it can be seen that the performance of the modelnaive bayes better than the modelknn anddecision tree. classification accuracy cannot achieve perfect results because there must be error values. this is influenced by the amount of test data and training data used in the simulation process carried out. 6. evaluation results with roc curve manual accuracy values can be done by looking at the roc curve comparison visualized from the confusion matrix. model viewing roc curves are the most easily visible way to graphically compare the accuracy values of each classification model. the graphical results of the roc can be seen in figures 10 and 11. figure 10 shows that the roc analysis results for student graduation at muhammadiyah university of pringsewu are correct in each model as follows: (1)k-nn is 0.500, (2) naive bayes is 0.500, and (3) decision tree is 0.600. therefore, for this case study, the model that has the best accuracy value isnaive bayes andk-nn because the curve approaches the point 0.1. figure 10. roc analysis with the true student graduation target figure 11 shows that the results of the roc analysis of late graduation of pringsewu muhammadiyah university students in each classification model are as follows: (1)k-nn is 0.500, (2) naive bayes is 0.500, and (3) decision tree is 0.600. therefore, classification research using 3 models with a study from muhammadiyah university of pringsewu is highly recommended using modelsnaive bayes andk-nn because the curve approaches the point 0.1. vol. 5, no.1 | 37 figure 11. roc analysis with the target of late student graduation based on the test results above, in this research the decision tree method has a slightly higher level of accuracy compared to the naive bayes method. there are several analyzes that cause the decision tree method to have higher accuracy, including the following: a. decision trees can work well on fairly large datasets without requiring excessive computing time. if the dataset is large enough, decision tree may be more efficient then decision tree provides an easy to interpret decision tree structure, which can assist in understanding the factors that most influence student graduation. b. naive bayes assumes feature independence, and normality of distribution. if these assumptions are not fully met in the dataset, the performance of naive bayes can be affected. decision trees, in some cases, are more resistant to this assumption. future research in the context of predicting student graduation on time using data mining methods can explore a number of aspects to improve the accuracy and sustainability of the model, including exploring the use of deep learning methods such as neural networks to understand more complex and non-linear patterns in the data and then considering factors time in analysis, such as changes in student behavior over time, curriculum changes, or changing campus policies and can develop algorithms that can provide better explanations for model decisions, especially in the context of academic decisions that can have major implications. iv. conclusions the resuts of this research show that after using the k-nearest neighbor, decision tree and naive bayes models to classify the graduation status of students at muhammadiyah university of pringsewu, the results obtained were that decision tree's performance was superior to k-nearest neighbor and naive bayes. it is proven that the data used by naive bayes has an accuracy value of 77.8%, a precision of 77%, while k-nearest neighbor has an accuracy value of 75.7%, a precision of 74% and the decision tree has an accuracy value of 77.9% and a precision of 78%. the contribution of this research can be used by the management of muhammadiyah university of pringsewu to detect early the condition of students so that their graduation is not too late and affect the accreditation score of muhammadiyah university of pringsewu. vol. 5, no.1 | 38 references [1] alim, s. (2021a). implementasi orange data mining untuk klasifikasi kelulusan mahasiswa dengan model k-nearest neighbor, decision tree serta naive bayes orange data mining implementation for student graduation classification using k-nearest neighbor, decision tree and naive bayes models. in jurnal ilmiah nero (vol. 6, issue 2). [2] amra, i. a. a., & maghari, a. y. a. (2017). students performance prediction using knn and naïve bayesian. icit 2017 8th international conference on information technology, proceedings, 909–913. doi: 10.1109/icitech.2017.8079967 [3] annur, h. (2018). klasifikasi masyarakat miskin menggunakan metode naïve bayes. in agustus (vol. 10, issue 2). [4] berrar, d. (2019). bayes’ theorem and naive bayes classifier. in s. ranganathan, m. gribskov, k. nakai, & c. schönbach (eds.), encyclopedia of bioinformatics and computational biology (pp. 403–412). oxford: academic press. doi: https://doi.org/10.1016/b978-0-12-8096338.20473-1 [5] caelen, o. (2017). a bayesian interpretation of the confusion matrix. [6] eko prasetiyo rohmawan. (2018). prediksi kelulusan mahasiswa tepat waktu menggunakan metode desicion tree dan artificial neural networ. [7] forsyth, d. (2018). probability and statistics for computer science. [8] hafizan, h., & putri, a. n. (2020). penerapan metode klasifikasi decision tree pada status gizi balita di kabupaten simalungun (vol. 1, issue 2). [9] kartini, d., nugroho, r. a., & faisal, m. r. (2017). klasifikasi kelulusan mahasiswa menggunakan algoritma learning vector quantization. in jurnal positif (vol. 3, issue 2). [10] mikut, r., & reischl, m. (2011). data mining tools. wires data mining and knowledge discovery, 1(5), 431–443. doi: https://doi.org/10.1002/widm.24 [11] parteek bhatia. (2019). data mining and data warehousing. [12] sari dewi. (2016). komparasi 5 metode algoritma klasifikasi data mining pada prediksi keberhasilan pemasaran produk layanan perbankan. [13] seref, b., & bostanci, e. (2018). sentiment analysis using naive bayes and complement naive bayes classifier algorithms on hadoop framework. 2018 2nd international symposium on multidisciplinary studies and innovative technologies (ismsit), 1–7. doi: 10.1109/ismsit.2018.8567243 [14] wati, e. f., & rudianto, b. (2022). universitas bina sarana informatika 1 teknik informatika, universitas nusa mandiri ,2 jl. kramat raya no.98, senen, jakarta pusat 10450 1 jl. in raya jatiwaringin (vol. 11, issue 2). teknik dan informatika. [15] wojtek j. krzanowski, d. j. h. (2009). roc curves for continuous data. [16] kartini, k., sujanto, b., & mukhtar, m. (2017). the influence of organizational climate, transformational leadership, and work motivation on teacher job performance. ijhcm (international journal of human capital management), 1(01), 192-205. [17] galit shmueli, p. c. b. i. y. n. r. p. k. c. l. jr. (2018). data mining for business analytics. 63 | vol. 3 no. 2 , july 202 2 p-issn : 2715-2448 | e-isssn : 2715-7199 vol.3 no.2 july 2022 buana information tchnology and computer sciences (bit and cs) a time series based gene expression profiling algorithm for stomach cancer diagnosis teresa kwamboka abuya1 study program computer science kisii university, kenya tkwambokaa@gmail.com ‹β› bayu priyatna 2 study program information system universitas buana perjuangan karawang bayu.priyatna@ubpkarawang.ac.id abstrak— eksperimen biologis telah menghasilkan sejumlah besar data ekspresi gen yang memiliki nilai sangat besar untuk diagnosis, pengobatan, dan pencegahan penyakit. namun, kelemahan yang cukup besar memang ada dalam pemanfaatan yang tepat dari data ini karena skala yang besar dan kerumitannya. sejumlah algoritma telah dikembangkan untuk menginterpretasikan data ini dalam bentuk profil gen untuk tujuan diagnosis. diantaranya k-means, pengelompokan hierarkis, pengelompokan berbasis kepadatan, pengelompokan subruang, dan peta yang mengatur sendiri. sayangnya, algoritme ini mengabaikan ketergantungan berurutan di antara titik waktu yang berurutan, tidak memadai dalam penemuan pola untuk mengubah aktivitas selama interval terbatas dari kerangka waktu eksperimen, dan tidak mampu membedakan antara pola faktual dan acak. dengan demikian, ada kebutuhan untuk algoritme pembuatan profil gen yang mengatasi kekurangan yang dibatasi waktu dalam algoritme saat ini dan karenanya memfasilitasi pembuatan profil gen yang efisien untuk mendiagnosis kanker perut secara dini. selama bertahun-tahun, eksperimen ekspresi gen deret waktu telah banyak digunakan untuk mempelajari berbagai proses biologis seperti siklus sel, perkembangan, dan respons imun. dalam makalah ini dikembangkan algoritma profil gen berdasarkan deret waktu untuk diagnosis awal kanker lambung. dengan menetapkan gen ke satu set profil model yang telah ditentukan sebelumnya yang menangkap pola potensial yang berbeda, signifikansi masing-masing profil ini dapat ditetapkan. profil signifikan ini kemudian dapat dianalisis lebih lanjut dan digabungkan untuk membentuk cluster yang kemudian dapat dimanipulasi oleh algoritma clustering. idenya adalah untuk mengukur aktivitas gen selama rentang waktu yang singkat sehingga dapat menghasilkan gambaran universal tentang fungsi seluler. singkatnya, ini termasuk mendeteksi pola berulang dalam data biologis. pola-pola ini kemudian digunakan untuk mengungkapkan informasi diagnostik yang mungkin penting bagi praktisi medis. desain penelitian eksperimental digunakan untuk mencapai tujuan penelitian. data yang berkaitan dengan genom biologis digunakan untuk pekerjaan penelitian ini. karena perkembangan penyakit kanker saat ini, hasil dari penelitian ini diharapkan dapat menjadi signifikan dalam diagnosis dini kanker lambung sehingga pengobatan yang tepat dapat diberikan.. kata kunci: data microarray, respon imun, clustering, profil signifikan, diagnosis kanker. abstract— biological experiments have produced enormous amount of gene expression data that possess enormous value for the diagnosis, treatment, and prevention of diseases. however, considerable drawbacks do exist in the appropriate utilization of this data due to its massive scale and intricacy. a number of algorithms have been developed to interpret this data in form of gene profiling for diagnosis purposes. they include k-means, hierarchical clustering, density-based clustering, subspace clustering, and self-organizing maps. unfortunately, these algorithms ignore the sequential dependency among successive time points, are inadequate in the discovery of patterns for changing activity over a restricted interval of an experiment’s time frame, and are incapable of discriminating between factual and random patterns. as such, there is a need for a gene profiling algorithm that addresses the time-constrained shortcomings in the current algorithms and hence facilitating efficient profiling of genes for early stomach cancer diagnosis. over the years, time series gene expression experiments have been widely used to study a range of biological processes such as the cell cycle, development, and immune response. in this paper a gene profiling algorithm based on time series for early stomach cancer diagnosis is developed. by assigning genes to a predefined set of model profiles that capture the potential distinct patterns, the significance of each of these profiles can be established. these significant profiles can then be analyzed further and combined to form clusters that can then be manipulated by clustering algorithms. the idea is to measure the genes’ activities over a short period span so as to come up with a universal depiction of the cellular functionality. in a nutshell, this includes detecting recurring patterns in biological data. these patterns are then employed to reveal diagnostic information that may be important for the medical practitioners. an experimental research design was utilized to achieve the study objectives. data pertaining to biological genomes was employed for this research work. due to the upsurge of cancer in the current times, the outcomes of this research work is anticipated to be significant in the early diagnosis of stomach cancer so that appropriate medication can be administered. keywords: microarray data, immune response, clustering, significant profiles, cancer diagnosis. i. introduction 64 | vol. 3 no. 2 , july 202 2 functional genomics is the discipline in which genes are utilized in the determination of their function whereas gene expression is an approach employed to examine the functional changes in these genes. according to [1], the expression level for a given gene across different experimental conditions are collectively referred to as the gene expression profile and the expression levels for all the genes under an experimental condition are jointly referred to as the sample expression profile. one of the goals in microarray data analysis is the identification of genes for which the expression level is significantly changed under different experimental conditions. another objective is to cluster the expressed genes or samples having similar expression profiles to make a meaningful biological inference from the set of genes or samples (martin et.al., 2016). the field of bioinformatics essentially deals with biological information processing. one of the requirements for effective bioinformatics is an extensive range of computational models that helps in representation and computation of massive biological data. as [2] point out, biological experiments and processes analysis require too much effort. additionally, this process can prove to be very slow. this can be attributed to the ever-growing intricacy of the processes and fiery growth of biological data emerging from laboratories universally. the recent drawback, as [3] noted, is on how to convert this enormous data repository into knowledge that can facilitate understanding of biological processes and experiments pertaining to both health and diseases. according to [4] timeseries gene expression analysis allows for principled estimation of unobserved time-points, clustering, and dataset alignment. in this technique, every expression profile is modeled as a piecewise polynomial which is estimated from the observed data and every time point sways the overall smooth expression curve. gene expression experiments carried out using time series show that unobserved timepoints can be reconstructed with 10-15% less error when compared to other profiling methods. the time series-based clustering algorithm operates directly on the continuous representations of gene expression profiles. this is particularly effective when applied to non-uniformly sampled data. stomach cancer (sc) is the fourth most frequently diagnosed malignancy and the second leading cause of cancer death worldwide (yang et.al.,2018). although the incidence of sc has declined for decades, the prognosis of sc remains very poor, especially in china. at present, the pathogenesis of sc is unclear, thereby necessitating effective biomarkers and targeted therapeutics. traditionally, clinic pathological parameters were used in risk stratification of sc outcomes. however, a number of advanced sc patients remained stable for a couple of years, whereas some early-stage patients progressed rapidly [5]. therefore, reliable biomarkers or stratification systems that can be used for more accurate prediction are highly essential [6]. the greatest challenge in cancer diagnosis is the identification of a subset of genes with crucial roles in diverse stages of these diseases’ progression from early stages of carcinogenesis to its final stage of metastasis. as [7] explains, reliable identification of molecular determinants of clinical outcomes can facilitate the discovery of functional biomarkers predictive of therapy response or disease progression. in addition, this can provide insights into new therapeutic targets in this aggressive disease. [8] further point out that the complexity of genomic networks and the vast volume of genes present increase the challenges of understanding and interpreting the resulting mass of data. the problem is compounded by the vagueness, imprecision, and noise present in this data. according to [9], the current algorithms such as hierarchical gene profiling algorithm, selforganizing maps (som), support vector machines (svm) and k-means algorithm, can only detect relationships where there is sufficient variability in gene expressions and as such, functional interactions are only detectable if they induce changes in transcriptional state that persist over a reasonable timescale. to address this problem, algorithms for visualizing high-throughput single-cell datasets and identifying putative functional relationships between genes are required [10]. due to the potential of time series to unravel biological processes that take place over short time duration, this research work employed this nonconventional data type to come up with a gene profiling algorithm that is instrumental in disease diagnosis in human beings. in this paper, a time series based gene profiling algorithm for early stomach cancer diagnosis was developed. early and accurate diagnosis of stomach cancer can significantly improve the design of personalized therapy and enhance the success of therapeutic interventions. since time series has the potential of identifying significant chronological expression profiles and the genes associated with this profile, it can enable the comparison of cancer infected genes behavior across multiple conditions over short time duration. specifically, the response of gastric epithelial cells infected with the vacamutant strain of the pathogen helicobacter pylori was investigated [11]. the contributions of this paper include the derivation of mathematical parameters that were shown to help in the generations of gene profiles over a limited duration of time. the rest of this paper is organized as follows. section 2 presents the related work while section 3discusses gene profiles derivation. section 4 gives a presentation of results and discussion while part 5 concludes the paper. ii. method this paper adopted an experimental research design to develop an algorithm that aided in the derivation of gene profiles. the approach involved the derivation of gene profiling parameters which were then employed to develop a time-series based algorithm. this algorithm was then experimented on sample genomic data described in section a below, to provide the required gene profiles visualization in the form of graphs. this visualization provided a straight forward means of establishing the sequential dependency among successive time point. in addition, the visualization facilitated the discovery of patterns for changing activity over a restricted interval of an experiment’s time frame. 65 | vol. 3 no. 2 , july 202 2 3.1 data set the genomics data employed in this paper were from two experiments measuring the response of gastric epithelial cells infected with the vaca-mutant strain of the pathogen helicobacter pylori. the data is sampled at five time points 0 h, .5 h, 3 h, 6 h, and 12 h. a sample of these data is shown in figure 1 for g27 tc1 trial 4. figure 1. sample g27 tc1 trial 4 data this figure 1 shows tc1 gastric epithelial (ags) cells infected with wild type h. pylori (g27) and isogenic mutants in caga and vaca for 0, 0.5, 3, 6, and 12 hours. figure 2.0 shows the g27 tc1 trial 5 data. figure 2. sample g27 tc1 trial 5 data in these data samples, hybridizations of g27 (trial 4) and cag a(trial 3) time-courses are accomplished in parallel. a technical replicate of the g27 time course (trial 5) and hybridization of vaca(trial 3) time course is also accomplished in parallel. the cag a 6and 12-hour time points technically replicated (trial 4) (the cag a 6-hour sample of trial 3 are lost). 3.2 gene profiling modeling process this research dealt with the profiling, comparing and visualizing gene expression data from short time series of two experiments measuring the response of gastric epithelial cells infected with the vaca-mutant strain of the pathogen helicobacter pylori. the gene expression profiling comprised of four major steps as shown in figure 3. as show in this figure, the steps included the generation and normalization of expression signals, testing each probe for its differential or association with the phenotype, the application of proper statistical significance criteria to identify the gene expression profile, and the investigation of the functions and pathways of the genes in the expression profile. figure 3. gene profiling steps thereafter, a number of statistical significance criteria such as pearson correlation, p-value, euclidean distance, logistic regression, bonferroni correction, false discovery rate and time points permutation were applied to help identify specific list of genes differentially expressed or associated with the phenotype. although mutual information (mi) measure is superior over simpler measures such as pearson correlation as it is capable of capturing complex non-linear and nonmonotonic dependencies. in addition, it can reflect the dynamics between pairs or groups of genes, computing mi involves estimating pair-wise joint probability distributions which requires density estimation or data discretization, with the accuracy of these estimates depending on sample sizes. as this measure was not deployed in this research study. table 2 gives a summary for the deployment of the various performance metrics. table 2 performance metrics deployment sno statistical measure deployment 1. pearson correlation weighted relation between all genes 2. p-value significance of gene coexpressions 3. euclidean distance correlation distance between gene profiles 4. logistic regression estimate of cancerous probability 5. bonferroni correction adjustment to the confidence levels 66 | vol. 3 no. 2 , july 202 2 6. false discovery rate adjustment to the confidence levels 7. time points permutation optimize the number of required profiles 3.3 the algorithm of modeling gene profiles the first step was the commencement of the algorithm while the second step in the processing activities was the input of the genomics data as shown in figure 4.0. in the next page. in step three, validation is done against empty genomic file upload such that if this field is empty, then an error message is generated for this effect in the fourth step. during the fifth step, the validation against spot ids not included is done such that if these ids are not included, then they are computed in the sixth step. the value of the spot id is initialized to 1 which are thereafter incremented by one until the value of 24192 is reached, which is equivalent to the number of genes in the file that were investigated. whereas spot ids were unique for each gene entry, the same gene symbol may appear multiple times in the data file corresponding to the same gene appearing on multiple spots. the seventh step was the computation of the average value for the expression values for the same gene. this was accomplished using the median before further analysis on the data was carried out. the eighth step was that option of filtering some specific genes using p-value metric. in situations where a gene was filtered, then it was excluded from further analysis. gene filtering was accomplished for those genes that did not show a sufficient response to experimental conditions, those genes that had too many missing values, or the gene expression pattern over repeats was too inconsistent as dictated by the minimum correlation between repeats. the ninth step was the usage of additional parameters namely the maximum pearson correlation and maximum number of candidate model profiles to dictate the selection of model profiles along with the maximum number of model profiles and maximum unit change in model profiles between time points as shown in figure 5.0. in this algorithm, the candidate model profiles were designed to be nonconstant profiles which started at zero and increased or decreased an integral number of units that was less than or equal to the value of the maximum unit change in model profiles between time points. figure.4. gene profiling algorithm pseudo-code when this parameter was set to zero, all permutations were used. in the eleventh step, the p-value based significance level was utilized to set the connotation level at which the number of genes assigned to a model profile as compared to the expected number of genes assigned was regarded as significant. during the twelfth step, the permutation test was set to permute all time points including time point zero when computing the expected number of genes assigned to a profile. in this case, the developed algorithm located profiles with significantly more genes assigned than expected on condition that all the input columns had been randomly reordered. on the other hand, during the thirteenth step, the permutation test was configured not to permute at time point zero and as such, the algorithm found profiles with more genes assigned than expected on condition that all the columns except for the first column had been randomly reordered. in the developed algorithm, permuting time point zero was preferred since it was the only test that took into account the significant changes that took place between time point zero and the immediate next time point (0.5 h). 67 | vol. 3 no. 2 , july 202 2 in the fourteenth step, the correction method was utilized to adjust the significance level since this algorithm was meant to test multiple profiles for significance. two types of corrections were utilized in this algorithm. the first one was the bonferroni correction while the second one was the conservative false discovery rate (fdr) control. in the third scenario, no correction was made for the multiple significance tests. figure 5. modeling gene profiling process in the fifteenth step, two parameters namely the minimum correlation and the minimum correlation percentile were utilized to control the grouping of significant model profiles into clusters. in so doing, these parameters served to control how similar two model profiles had to be if they were grouped together. for the case of the minimum correlation, any two model profiles assigned to the same cluster of profiles had to have a correlation above this parameter's value. on its part, the minimum correlation percentile was employed in cases there were repeat data from different time periods. it was used to specify that any two model profiles assigned to the same cluster of profiles had to have a correlation in their expression greater than the correlation of this percentile in the distribution of gene expression correlations between the repeats. the last step was the display of the gene profiles based on the euclidean distance after which the algorithm halted in the seventeenth step. figure 6 gives a diagrammatic representation of the gene profile derivation process. as this figure shows, the process gene profile derivation process involves the input of the genomic data containing the gene expressions to be profiled. these data items are analyzed using parameters such as pvalue, pearson correlations, permutations, logistic regression and median to yield probable profiles as already discussed above. correction methods are then employed to adjust the significance level to permit the testing of multiple profiles for significance. figure 6. schematic gene derivation process the output gene groupings are then clustered using minimum correlation and minimum correlation percentile before euclidean distance is applied to them to distinguish the various gene profiles. the final outputs are the gene profiles in form of graphs. the logic here was that when the number of candidate model profiles exceeded the p value of seeing t more genes in the intersection, then instead of explicitly generating al l candidate model profiles, a subset of candidate model profiles of this size was randomly selected. in the tenth step, the number of permutations per gene parameter was employed to specify the number of permutations of time points that were randomly selec ted for each gene when computing the expected number of genes assigned to each of the model profiles. 68 | vol. 3 no. 2 , july 202 2 iii.results and discussion in this section a time series-based gene profiling algorithm is developed. to test the derived parameters and their gene profiling abilities, the algorithms and statistical computations were put into use to achieve some functionality as shown in table 3. the genomics data from two experiments measuring the response of gastric epithelial cells infected with the vac a mutant strain of the pathogen helicobacter pylori were then fed as input to this algorithm. table 3. gene derivation process step parameter activity 1 n/a -commence gene derivation process 2 n/a -input genomic data 3 n/a -validation is done against empty genomic file upload 4 n/a -prompt genomic data input error 5 n/a -validation against spot ids 6 n/a -if not included in file compute spot ids 7 median -computation of the average value for the expression values for the same gene 8 p-value -filtering specific genes 9 pearson correlation, p-value -model profiles selections. 10 permutation -computation of the expected number of genes assigned to each of the model profiles 11 p-value, logistic regression -setting the connotation level at which the number of genes are assigned to a model profile 12 permutation -compute the expected number of genes assigned to a specific profile 13 permutation -configure permutation test not to permute at time point zero 14 bonferroni, fdr -adjust the significance level to test multiple profiles for significance 15 minimum correlation, minimum correlation -control grouping of significant model profiles into clusters percentile 16 euclidean distance -display generated gene profiles 17 n/a -halt gene derivation process the minimum absolute expression change was any value more than -0.05. as an illustration, using the maximum number of missing values to be 2, the minimum correlation between repeats to be 0, and the minimum absolute expression change to be 0.05 yielded the information in table 4.0 for the sample filtered genes. table 4. sample filtered genes the genes that were devoid of these three characteristics were regarded as standard genes and were the ones that took part in further analysis. table 5 gives information on the sample genes that passed the classification criteria. table 5. sample genes passing classification criteria afterwards, eight parameters were utilized for the computational derivation of gene profiles from this set of data: maximum correlation, maximum number of candidate model profiles, maximum number of model profiles and maximum unit change in model profiles between time points, number of permutations per gene, significance level, and correction method as shown in table 6 below. table 6. gene profiles evaluation metrics gene profiling option value maximum correlation 1 maximum number of candidate model profiles 1,000,000 number of permutations per gene(0 for all permutations) 0 p-value significance level 0.05 maximum number of model profiles 50 maximum unit change in model profiles between time points 2 correction method none minimum correlation 0.7 based on the evaluation metrics of table 6.0, the algorithm was run to yield the proposed gene profiles. 3.1 time series-based gene profiling. in this profiling, the maximum correlation specified the value that the maximum correlation between any two model profiles had to be below, and was therefore employed to guarantee that two very similar profiles were not selected. the maximum value for this parameter was set to 1 in order to prevent two perfectly correlated model profiles from being selected. it was observed that lowering this parameter led to the number of model profiles selected being less than the maximum number of model profiles even in situations where more candidate model profiles were available. on the other hand, the maximum number of candidate model profiles represented non-constant profiles which commenced at 0 and increased or decreased an integral number of units that was less than or equal to the value of the maximum unit change in model profiles between time points. the number of permutations per gene parameter specified the number of permutations of time points that were randomly selected for each gene when computing the expected number of genes assigned to each of the model profiles. 69 | vol. 3 no. 2 , july 202 2 when this parameter was set to 0, all permutations were used. it was also important to set permutation test to permute time point 0 or not. when computing the expected number of genes assigned to a profile, if the permutation test for time point 0 was set, the permutation test permuted all time points including time point 0. it was observed that doing this led to profiles with significantly more genes being assigned than expected if all the input columns had been randomly reordered. on the contrary, if the permutation test was not set for 0, the permutation test permuted all time points except for time point 0. in this scenario, profiles with more genes were assigned than expected if all the columns except for the first column had been randomly reordered. permuting time point 0 was preferred since only this test took into account significant changes that occurred between time point 0 and the immediate next time point. however in some cases based on experimental design a gene's expression value before transformation at time point 0 was expected to be known more accurately than the other time points, and because of this asymmetry, not permuting time point 0 was also be useful. it was observed that increasing the maximum number of model profiles increased the number of candidate models as shown in table 7. table 7. maximum number of gene model profiles viz. significant gene models maximum number of model profiles resulting sig. number of gene models 50 14 60 16 70 18 80 20 100 21 120 23 140 24 based on the values in table 7, a graph was plotted for maximum number of model profiles against the resulting significant number of gene models as shown in figure 7. figure 7. maximum no. of model profiles viz. resulting significant. no. of gene models the graph of figure 7.0 shows that the resulting significant number of gene models increase nearly exponentially as the maximum number of model profiles was increased. consequently, to get fine grained gene model profiles, the maximum number of model profiles had to be increased and vice versa. table 8.0 gives the shift in the resulting significant number of gene profiles as the maximum unit change in model profiles between time points was adjusted. table 8. maximum unit change in model profiles viz. significant gene models maximum unit change in model profiles resulting sig. number of gene models 1 13 2 14 3 15 4 13 5 14 6 15 7 15 8 16 9 16 10 15 as shown in this table, generally as the maximum unit change in model profiles is increased, the resulting significant number of gene models is increased. this is due to the pronounced euclidean distances between the gene models. regarding maximum correlation, the value of minimum correlation was set to zero (0) and the value of maximum correlation was slowly reduced from 1 to zero. the results obtained are shown in table 9 are observed. table 9. maximum correlations viz. significant gene models maximum correlation resulting sig. number of gene models number of genes assigned 1 12 1005 0.9 12 1005 0.8 9 993 0.7 6 1185 0.6 4 1216 0.5 3 1243 0.4 3 1348 0.3 3 1348 0.2 3 1470 0.1 3 1470 0 1 1275 generally, as the value of maximum correlation is reduced from one to zero, the number of resulting significant number of gene models reduced to unity (1) while the number of genes assigned to these gene models increased from 1005 to a maximum value of 1275. this implies that when the correlation value is small, gene models are basically indistinguishable hence at correlation zero, there is only one resulting significant model. on the other hand, at maximum correlation, the genes can be clearly distinguished and hence the resulting significant gene models are many. concerning the number of genes assigned to models, at low correlation coefficients, genes profiles are indistinguishable and hence a large number of genes are assigned to the few available models. however, as the correlation coefficients are increased, the gene profiles become increasing disparate and 70 | vol. 3 no. 2 , july 202 2 few genes are assigned to each of the many models now available as the rest are discriminated due to their large pvalues. 3.2 prediction power of the developed algorithm in the developed algorithm, sequences of gene expressions were listed in order of occurrence, starting at time point 0h to 12h. the aim was to collect and investigate precedent observations of gene expressions at various time points in order to come up with ideal models to express the intrinsic structure of the underlying genomic data. based on these models, it was possible to predict future gene expressions. to put this into perspective, profile id 17 was considered whose gene expressions are shown in figure 8. figure 8. gene expressions for profile id 17 a total of 54 genes were assigned to this model profile whose individual expressions are shown in figure 8. by sketching a line of best fit through these gene expressions and performing some extrapolations, the future expressions beyond the 12h time point can be obtained as shown in figure 9 below. figure 9. gene profile prediction the white thick line through the gene expressions is the line of best fit while the thick red line represents the extrapolated gene expressions for the 54 genes assigned to profile id 17 for the future 18h and 30h time points. suppose that the stomach cancer patient gene expressions are as shown in figure 10 below. figure 10. stomach cancer patient gene expressions over 30h duration comparing the hypothesized gene expressions over the 30h duration and the gene models in figure 11.0 below, then considering the first few gene expressions, model profile ids 13, 14,15,16,17 and 18 are candidates’ models that the stomach cancer patient gene expressions can fit in. however, taking into account the preceding time points eliminates model profiles 14(experiences near exponential growth followed by plateau), 15(experiences linear growth followed by plateau), 16 (portrays linear growth, linear decay and plateau), and 18 (presents linear growth followed by plateau). this leaves profile id 13 and 17 as the most probable model profiles. by drawing a horizontal line through these two profiles as shown in figure 11, it is possible to discern which of them perfectly fits the patient gene expressions. figure 11. gene model fitting based on this line and considering time-points at which troughs and crests appear, it is clear that model profile id 17 perfectly fits the patient gene expressions for a duration of 30h time points. as such, it can be implied that the developed algorithm led to accurate diagnosis of stomach cancer patients within 12h time points since the commencement of the cancerous gene expressions. in the next section, this algorithm is validated against some well-known gene profiling algorithms. 3.3 validation of the developed algorithm in this section, the time series-based algorithm that was developed is validated 71 | vol. 3 no. 2 , july 202 2 against other gene profiling algorithms such as hierarchical gene profiling algorithm, support vector machine, selforganizing maps, and k-means algorithm. in hierarchical gene profiling algorithm, genes with related expression patterns are grouped together and connected by a series of branches to form a dendrogram. unfortunately, this algorithm considers each gene as an individual cluster and genes that are similar to each other form nested clusters based on the pair-wise distances. on the other hand, the time series-based algorithm developed in this research study considered a group of genes with similar expressions as profile clusters. for instance, in figure 6.6, a total of 155 genes were represented by a single model profile with id 40 and 90 genes were represented by model profile id 37. these two model profiles formed a cluster with a total of 245 genes. as such, the developed algorithm is operationally faster during gene profiling compared to hierarchical gene profiling algorithm, rendering it ideal for large genomic data set. the genomic data that was utilized in this research consisted of 24192 gene symbols observed under 5 time points, making the total gene expressions 120960, a very big data set for the rather slow hierarchical gene profiling algorithm. to effectively apply the support vector machine gene profiling algorithm, it requires training using the same members of each model profile that have to be identified. this training takes time and hence compared to the developed algorithm, it is slow and hence inefficient for large data sets such as the 120960 gene expressions that were under investigation in this research. although self-organizing maps algorithm has been employed to group 1,036 genes into 24 categories, this algorithm is slow in training, hard to train against slowly evolving data and are not so intuitive since neurons close on the map (topological proximity) may be far away in feature space. additionally, these maps do not behave so gently when using categorical data, or mixed data. comparing the 1036 genes that selforganizing maps algorithm profiled into 24 categories with the 24192 genes that were profiled using the developed time series-based algorithm, it is clear that the proposed algorithm is efficient. regarding svm algorithm, this algorithm has been used for cancer classification with microarray data where it served as a powerful classifier together with four effective feature reduction methods namely principal components analysis (pca), class-separability measure, fisher ratio and t-test to the problem of cancer classification based on gene expression data. although it very high classification accuracies, it requires feature reduction methods which renders it structurally complex compared to the time series-based algorithm implemented in this research. on its part, the k-means algorithm operates on a series of microarray experiments measuring the expression of a set of genes at regular time intervals in a common cell line. it requires that data be normalized to permit for comparisons across these microarrays. the output produced is in form of clusters of genes which vary in similar ways over time and hence it is possible to infer that genes which vary in the same way may be co-regulated and or participate in the same pathway. unfortunately, the numbers of clusters need to be specified which may be unknown in some instances, and figuring out the right number of clusters that represent the true number of clusters in the population is quite subjective. as such, the profiles obtained using k-means can vary greatly depending on the location of the observations that are randomly chosen as initial centroids. however, the developed time series-based algorithm employs statistical metrics such as p-value, pearson correlation, logistics regression and euclidean distance whose significance levels are well known. the k-means clustering algorithm assumes that the underlying clusters in the population are spherical, distinct, and are of approximately equal size and hence tends to identify clusters with these characteristics. therefore, this algorithm is incapable of yielding good results when clusters are elongated or not equal in size like the genomic data used in this research where some gene expressions were negative, zero and others positive. the kalgorithm is also sensitive to initial conditions, implying that different initial conditions produce varying result of gene profiles. it is also possible for a very far data from the centroid to pull the centroid away from the real one as shown in figure 12. below. here, 5 genes are assigned to cluster id 55 and it is clear from the gene. figure 12. k-means based profiling expressions that the profiling is not such accurate especially after the 0.5h time point. whereas 4 gene expressions have negative gradients, one of them has a positive gradient. during the 3h time-point, some gene profiles are at the rough, others are at the crest, plateau while others are still on their descent. the same is observed during the 6h time point. these varying result of gene profiles give contradicting depiction of gene activities and hence may lead to inaccurate stomach cancer diagnosis. iv.conclusion the aim of this paper was to develop a gene profiling algorithm based on time series to help in early stomach cancer diagnosis. based on a number of derived gene profiling 72 | vol. 3 no. 2 , july 202 2 parameters, an algorithm was developed that was then experimented on sample genomic data. the results of this paper included a number of gene profiles that obtained from the underlying pathogen helicobacter pylori data. the significance of this research lies on the fact that it helped generate gene profiles using very short time points. this feature is very critical in early stomach cancer diagnosis as it facilitates necessary preventive measures that curtail the cancerous cells advancement to other fatal phases. since this research was purely based on stomach cancer, future work in this area lies on the implementation of this algorithm for other types of cancer or diseases. references [1] sanchita & ashok s. (2015). future challenges in application of algorithms and tools for clustering of gene expression data. biotechnology division, csircentral institute of medicinal and aromatic plants. lucknow 22601. (pp. 515-531)5 india. [2] brohée s., barriot r., & moreau y. (2015).biological knowledge bases using wikis: combining the flexibility of wikis with the structure of databases. bioinformatics, oxford journals. [3] wong k. (2016).computational biology and bioinformatics: gene regulation. crc press. [4] ziv b., georg g., david k., & tommi s. (2015). a new approach to analyzing gene expression time series data. whitehead institute for biomedical research. [5] wang, h., wang, x., xu, l. et al. (2020).high expression levels of pyrimidine metabolic rate–limiting enzymes are adverse prognostic factors in lung adenocarcinoma: a study based on the cancer genome atlas and gene expression omnibus datasets. purinergic signalling 16, 347–366 (2020). https://doi.org/10.1007/s11302-02009711-4. [6] wenhui, y., zhiyong l.., yuan, li., jianbing, m., mudan, yang., jun, x.(2019).immune signature profiling identified prognostic factors for gastric cancer. chinese journal of cancer research. https// doi: 10.21147/j.issn.1000-9604.2019.03.08.[7] abolfazl r., fatemeh a., salendra s., and vinay v. (2016).networkbased enriched gene subnetwork identification: a game-theoretic approach. biomed eng comput biol. vol. 7, issue 2, pp. 1–14. [8] oyelade j., itunuoluwa i., funke o., olufemi a., efosa u., faridah a., moses a., and ezekiel a. (2016). clustering algorithms: their application to gene expression data. bioinformatics and biology insights,10, 237–253. [9] thalia e.c., michael p.h., and ann c. (2017).gene regulatory network inference from single-cell data using multivariate information measures. cell systems, 5, 251–267. [10] gwang h., peter s., sung j., & joo h. (2016). screening and surveillance for gastric cancer in the united states: is it needed? american society for gastrointestinal endoscopy. volume 84, no. 1, pp. 18-28. [11] siregar, amril mutoi, et al. "perbandingan algoritme klasifikasi untuk prediksi cuaca." jurnal accounting information system (aims) 3.1 (2020): 15-24. vol. 5, no.2 june 2024 | 74 ransomware detection using machine learning algorithm olaniyi abiodun ayeni 1*, ibitola elizabeth adejumo 2 department of cyber security, school of computing, federal university of technology, akure nigeria* department of academic planning, ict, university of medical sciences, ondo state nigeria. e-mail: oaayeni@futa.edu.ng1*, iadejumo@unimed.edu.ng 2 abstract with the advent and subsequent explosion of the internet, global connectivity has been achieved, and is on the rise. this provides a host of advantages such as connectivity and communication, information broadcast and transmission, amongst others. this however introduces a new set of challenges: the safety and protection of these communication channels amongst them. information has always been power, and the widespread mature of information only results in the widespread attempts to procure it, sometimes via illegal channels. in view of this, this research aims at detecting crypto-ransomware and locker ransomware. data was collected from an open repository and cleaned. the cleaned data was then split into tests, train sets and validation which was used to train a number of ml models based on the: random forest algorithm, support vector machine (svm) and gradient boosting algorithm. ransomware is one of the well-known ways and frequent use which cyber-attackers use in infecting their victims, either through phishing or drive download. attackers will create an email pretending to be from a genuine resource and send it to their targeted victims. however, this research illustrated how to combat crypto-ransomware and locker ransomware. implementing the machine learning algorithm, the system can detect ransomware under 30’s, giving computer users over 90% assurance of their system for ransomware free. keywords: gradient boosting algorithm, machine learning, random forest, ransomware, support vector machine. i. introduction it is possible to access information via the internet and easily recover it for a cheaper cost in our digital world today where information’s are stored digitally. without stress, everything is completed effortlessly and efficiently. digitalization has increased computer users' quality of life. but every pillar has two sides, as the saying goes. if used as a whole, digitization has reduced crime since technology has made tasks easier to complete and requires less paperwork. however, it creates a security concern for a person's private and sensitive data information. there are numerous thefts and cyberattacks that have occurred which include viruses, spyware, malware, trojans, phishing, and intruders [1]. ransomware is a theft which is a kind of infection that can be hard to recover from when being spread. important files and data on the user's computer system are corrupted as a result of this ransomware. ransomware is a kind of malicious malware where the attackers encrypt your file and make it inaccessible to the owner, which spreads more widely and gets more sophisticated every day [10]. every system in the network today is susceptible to attacks by online criminals. now that automated technologies are more sophisticated, attackers have access to them, and new threats appear almost instantly. this makes it possibly challenging to maintain proper cybersecurity. malicious software is one of the biggest concerns in the digital world, and sadly, the problem is getting worse day by day [4]. p-issn: 2715-2448 | e-issn: 2715-7199 vol.5 no.2 june 2024 buana information technology and computer sciences (bit and cs) mailto:oaayeni@futa.edu.ng mailto:iadejumo@unimed.edu.ng vol. 5, no.2 june 2024 | 75 manabu et al. (2019) stated that with the quick expansion in internet of things (iot) devices, cyberphysical systems mobile devices, and the cloud services, there has been a surge in extensive cyberattacks on businesses and governmental sectors. specifically, ransomware is a type of software that prevents victims from accessing or making use of their systems and files until a ransom is paid. the current level of cyber security today is an ongoing process that entails gathering and comparing millions of data points across all of the personnel and infrastructure. it is fairly obvious that relying only on humans would not be sufficient, there is a need for machine learning support in order to identify trends and foresee potential security risks in massive data sets. ransomware can be classified into three categories. fig 1. types of ransomware crypto-ransomware is a type of ransomware that encrypts some vital files in the computer system making the user not able to access files or make use of the computer, for the user to retrieve the files in the system, then a ransom message will be passed to the computer user or the victims demanding payment before the user can retrieve the file or information back which can either be retrieved or not after payment and a short period of time will be given for the payment of the ransom. the attackers get their ransom by holding the vital files hostage and this ransom request is through a means like bitcoin. an example of this ransomware is wannacry [7]. crypto-ransomware can be further being sub-divided into three which are: a. symmetrical crypto-ransomware b. asymmetrical crypto-ransomware c. hybrid crypto-ransomware. this is the type of ransomware that infects the user’s computer and blocked the user out completely, preventing the user from accessing their files on the computer. even some parts of the computer can be blocked also the keyboard. the attacker thereafter demands a ransom to unblock the computer and limited access will be given to the user to communicate with the attacker until the ransom is paid, the computer will be unblocked. despite the disruption caused by ransomware, the computer user can still easily be retrieved by removing the disk from the compromised system and placing it on a well-cleaned system [8]. the locker ransomware process is as follows. fig 2. locker ransomware process vol. 5, no.2 june 2024 | 76 scareware is ransomware with a tricky technique or a false message used to fake computer users in order to convince users to download harmful software or ransomware that can encrypt data and demand payment. this kind of ransomware post no danger to the victim. attackers take advantage of the fear of the users to attack the victims [8]. ii. review of related works the work of [6], dynamic feature dataset for ransomware detection using machine learning algorithms aims to conduct some analyzing and selecting the most relevant and non-redundant dynamic features for identifying encryptor and locker ransomware from goodware, generating json files with dynamic parameters using a sandbox through experiments with encryptor and locker ransomware combined with goodware, and applying the dynamic feature dataset to obtain models with machine learning algorithms.. a dynamic features dataset is generated and made public. this method made use of machine learning methods, static and dynamic ransomware analysis, and a dataset derived from the created json files. at last, a dataset was created that included traits taken from decent software and the dynamic aspects of both locker and encryptor ransomware. however, a dataset was developed. a dataset was developed. the study's dataset contains relevant and lightly correlated features linked to ransomware that is created in runtime. eduardo et al. (2022), presented crypto-ransomware detection using machine learning models in file-sharing network scenarios with encrypted traffic. this focuses on offering an algorithm validation through an analysis of the false positive rate and the volume of user file data that the ransomware could encrypt before being discovered through deep learning and machine learning (neural network model optimization), as well as model validation through the use of various file-sharing protocols. while the malware is reading and writing files to a network-shared disk, the research finds crypto-ransomware. provide a tool for detecting crypto-ransomware that is based on the examination of encrypted network traffic in situations involving file sharing. to do this, capabilities for extracting and filtering data that can differentiate between ransomware activity and innocuous, high-activity traffic must be used. meanwhile, this system detects crypto-ransomware while the malware is reading and writing files in a network-shared volume with a high false positive. computers and mobile operating system were not considered. samah et al. (2019), presented ransomware detection system for android applications. this study suggested a static analysis mechanism for locating android ransomware programs. based on the calls made by api packages as a leading indicator of harmful behavior, api-rds focuses on identifying ransomware with high accuracy before it damages the user's device. analyze the most recent approaches to ransomware detection for android devices by collecting information, suggesting an api-based system (api-rds), evaluating api-rds, and finally providing api-rds services. this dataset includes information from apipackages calls, android ransomware, and innocuous android datasets. for the purpose of identifying android ransomware apps, the research offers a static analysis paradigm. it also creates a unique and current dataset that includes recent clean apps and most of the current android ransomware families. this labeled reference might be applied by the research community. however, the system only focuses on android application detection which may not be applicable to other application. sh kok et al. (2019), in a study titled ransomware, threat and detection techniques: a review this essay discusses the most recent methods for detection and offers a comprehensive overview of the threat posed by ransomware. provided a thorough description of the steps involved in a ransomware attack and their traits, which can be used as a foundation for further ransomware study. data from static and dynamic assessments were combined to create a hybrid algorithm for approach. but the it is a review work and no model for a ransomware attack has been developed. the researcher has suggested that in the next research, a model to identify ransomware attacks be developed and that hybrid algorithms be used in place of a single one. subash et al. (2019), proposed a multi-level ransomware detection framework using natural language processing and machine learning. the researchers presented a multi-level big data mining system that combines methodologies from machine learning, natural language processing (nlp), and vol. 5, no.2 june 2024 | 77 reverse engineering. detector engine, action engine, passive analyzer, function call tracker, assembly instruction tracker, dll tracker, and six other components are used in natural language using machine learning. the open source malware repository zoo and virus total were two of the sources from which the dataset was gathered. this study created a framework for multi-level analysis using dlls, function calls, and assembly instructions while taking advantage of machine learning classifiers and nlp schemes. it also investigated the differences in n-gram sequences for ransomware binary samples at the multi-level, which helped to create a useful feature database that increased the detection rate at various levels. 98.59% is the maximum detection accuracy for n-gram tf-idf at n=3, and 97.13% is the second-highest at n=2. nevertheless, the researcher admitted that performance testing between the research framework and nlp schemes and machine learning classifiers was not done. instead, a framework of multi-level analysis was built using dll function calls and assemble instructions. iii. method in order to improve the detecting performance of ransomware, the architecture of the suggested ransomware detection system was designed to improve detecting capabilities, the system incorporates machine learning techniques. the goal of this all-encompassing approach is to give enterprises a strong and flexible defense against the constantly changing threat landscape that ransomware attacks is present. fig 3. system architecture a. data collection the dataset of ransomware attack instances was obtained from kaggle.com in excel format. this dataset includes both benign and malicious samples, covering various types of ransomware. initializing threshold settings, outliers will be checked by comparing the distance of the closest data point to the nearest cluster identification and identifying those that are outliers in our dataset. the dataset represents real-world scenarios and contains features relevant to ransomware detection, such as file characteristics, network traffic patterns, and behavioral indicators. the dataset contains 143573 rows and 84 columns. vol. 5, no.2 june 2024 | 78 fig 4. sample of ransomware dataset b. data preprocessing the data were cleaned and preprocess the collected data to remove any noise or inconsistencies handling missing data and identify and handle the missing values or data. by removing rows or columns with missing values or data and imputing the missing values. outliers were identified and handled by removing them with the use of a robust statistical method. perform feature selection to extract meaningful features from the raw data. this may involve techniques like dimensionality reduction, feature selection, or transformation. the data was cleaned appropriately, after which particular features were selected. c. feature selection feature selection was carried out for ease of modeling. the data was truncated for ease and speed of modeling. perform required feature selection and dimensionality reduction. perform model selection based on the algorithm choices outline (i.e., random forest, svm and gradient boosting). feature selection is a crucial step in machine learning where the goal is to choose the most relevant and significant features from a dataset to build a model. the process involves identifying and selecting a subset of features that contribute the most to the model's predictive power while disregarding irrelevant or redundant ones. this is done to improve model efficiency, reduce overfitting, and enhance generalization to new. a. random forest an ensemble learning method for applications like categorization and regression is called random forest. additionally, during the phase, random forest constructs a large number of decision trees and produces a class that represents the mean of the classes, also known as classification, or mean prediction, also known as regression of each of the trees. it is common for random forests to correctly predict their training set. random forest is the go-to machine learning algorithm that uses a bagging approach to create a bunch of decision trees with a random subset of the data. a model is trained several times on a random sample of the dataset to achieve good prediction performance from the random forest algorithm. the output of every decision tree in the random forest is pooled to provide the final prediction in this ensemble learning technique. by analyzing the outcomes of each decision tree, the random forest method's final prediction is discovered or by selecting the forecast that emerges most frequently from the decision trees. the training set, the test set and validation are the three subsets that make up the random forest, and it selects some samples from the practice set [5]. random forest aims at lowering the amount of time needed for learning and classification either to seek to increase accuracy, performance or both. the random first extracts subsamples from the original samples with the aid of the bootstrap resampling technique, then the algorithm categorizes the decision trees and implements a simple vote with the classification’s largest vote serving as the prediction’s outcome. there are three steps in the random forest algorithm which are: vol. 5, no.2 june 2024 | 79 1. choose the training set by retrieving training sets from the original dataset employing the bootstrap random sampling technique, making sure that each training set has the same size as the first training set. 2. develop the random forest model by making a classification regression tree for every bootstrap training set. these trees are left untrimmed to generate decision trees that make up the forest. 3. create simple voting: decision trees that have been trained in the same manner can be combined to generate the random forest. because each decision tree's training procedure is autonomous, training for random forests can proceed simultaneously, greatly enhancing efficiency xiang et al. (2019). b. support vector machine (svm) it permits the search for nonlinear decision boundaries using a variety of various kernels, support vector machine can be used to categorize points from a data set in nonlinear decision boundary. support vector machine basis has four possible values which are sigmoid, linear, polynomial and radial which is called kernel parameter [3]. c. gradient boosting algorithm gradient boosting is the method that enables gradual construction of an ensemble trees with the aim of reducing a target loss function. boosting keeps the leaf node labels and the weights in a way that makes handling prediction interpretations simple. xgboost is one of the classification methods. the two improvements in xgboosting over gradient descent are its improved periodicity technique and its increased level of sophistication. gradient boosting retrieves the relative value scores of each attribute following that the boosted tree is built using an effective metric known as feature/importance [2]. d. model training and evaluation in this research, the model will be trained using 75% of the data, 15 for testing and the remaining 10% will be used for validation. four criteria will be used to evaluate the trained model’s performance on the testing set, considering metrics like precision, recall, f-score and accuracy. precision = 𝑇𝑃 𝑇𝑃+𝐹𝑃 (1) recall = 𝑇𝑃 𝑇𝑃+𝐹𝑁 (2) f-score = 2. (𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛∗𝑅𝑒𝑐𝑎𝑙𝑙) (𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛+𝑅𝑒𝑐𝑎𝑙𝑙) (3) accuracy = 𝑇𝑃+𝑇𝑁 𝑇𝑃+𝑇𝑁+𝐹𝑃+𝐹𝑁 (4) where: tp = true positive values which stand for values that are correctly anticipated to be positive. fp = false positive values which stand for values that are incorrectly expected to be positive. fn = false negative values which stand for values that are incorrectly predicted to be negative. tn = true negative values which stand for values that are correctly predicted to be negative. e. mode training hyperparameters such as `n_estimators`, `max_depth`, and grid search to optimize model performance were used for training the model, numpy (python library used for working with arrays). imports sklearn, numpy, pandas, imbalanced-learn, cuml, matplotlib. sets up rapids python env with cuda libraries for gpu acceleration. vol. 5, no.2 june 2024 | 80 iv. experiment result the experiment was on the three-model used which are: random forest, support vector machine on the dataset and are compare. the gradient boosting model looks most robust, maintaining high scores on the test set. the others overfit slightly more. the overall accuracy scores look quite good, all above 98-99% for train and validation. this suggests the models are fitting the training data very well. however, the test accuracy is a better evaluation of real-world performance. here the scores drop slightly but are still strong at 98-100%. precision and recall scores are also generally high, indicating the models are successfully learning the patterns in the data. the f1 scores are balanced, not favoring precision or recall heavily. the macro averages show there isn't a huge skew towards any particular class. the high accuracies and f1 scores indicate the models are learning the patterns and generalizing fairly well. some overfitting is present but performance remains strong. below is the comparison result for the models on training, testing and validation. table 1. showing training classification model accuracy recall precision f1 random forest svm gradient boosting 0.98 1.0 1.0 0.99 0.97 0.99 0.97 0.97 0.99 0.98 0.97 0.99 table 2. showing testing classification table 3. showing validation classification model accuracy recall precision f1 random forest svm gradient boosting 0.98 0.99 1.0 0.99 1.0 1.0 0.97 0.99 1.0 0.98 0.99 1.0 model accuracy recall precision f1 random forest svm gradient boosting 0.98 1.0 1.0 0.99 1.0 1.0 0.97 0.99 1.0 0.98 0.99 1.0 vol. 5, no.2 june 2024 | 81 fig 5. graphical illustration of training classification fig 6. graphical illustration of testing classification fig7. graphical illustration of validation classification vol. 5, no.2 june 2024 | 82 a. confusion matrix the selection of which metrics to prioritize is contingent upon the particular objectives and demands of the task at hand. these metrics offer distinct viewpoints on the model's performance. fig 8. confusion matrix for random forest fig 9. confusion matrix for svm fig10. confusion matrix for gradient boosting validation b. comparison on final performance on the model the final performances of the models were compared over the roc auc score. an evaluation was carried out via the roc auc metric which is a tool for assessing and comparing the performance of classification models, particularly in situations where the balance between false positives and false negatives is important. these are tabulated as follows for the datasets. vol. 5, no.2 june 2024 | 83 table 4. showing final performance for the models of the dataset model train auc 75% test auc 15% valid auc 10% generalization error (%) random forest svm gradient boosting 0.962 0.923 0.9998 0.899 0.767 0.9939 0.98 0.99 1.00 6.3 15.6 0.59 figu11. graphical illustration of the final performance fig 12. auc comparative analysis v. conclusion ransomware detection is an ongoing and multifaceted challenge that requires a combination of advanced technology, user awareness, and a proactive cybersecurity posture. private users, commercial enterprises, and government networks must invest in modern detection techniques. the use of machine learning in ransomware detection has great potential to improve cybersecurity defenses. nevertheless, by applying machine learning algorithms like gradient boosting, random forest, and support vector machines, this study has been able to offer advice on how to cope with both locker and crypto-ransomware. users of computers can be more than 90% confident that their system is clear of ransomware. after the models were vol. 5, no.2 june 2024 | 84 compared, it was found that the gradient boosting model (99%, 99%, and 100% auc train, test, and validation, respectively) had the lowest generalization error, while the svm model (92%, 76%, and 99% auc train, test, and validation) performed the poorest. despite this, the models were still overfit. in the middle of the pack (96%, 89%, and 98% auc) was the random forest model. hyperparametric optimization can be used to reduce the generalization error. the results, however, show that machine learning techniques have a lot of potential for use in cybersecurity and ransomware detection. references [1]. abdullahi arabo, remi dijoux,timothee poulain,gregoire chevailer, (2020), detecting ransomware using process behavior analysis. pp. 289 and 295 [2]. darshana u., jaume m., marzia z., and srinivas s. (2019). gradient boosting feature selection with machine learning classifiers for intrusion detection on power grids. ieee transactions on network and service management. pp. 3-5. [3]. drew conway and john myles white (2012) machine learning for hackers. first edition http://oreilly.com/catalog/errata.csp?isbn=9781449303716 o’reilly media, inc. pp. 275-278. [4]. eduardo berrueta, daniel morato, eduardo magana, mikel izal (2022), crypto-ransomware detection using machine learning models in file-sharing network scenarios with encrypted traffic. pp. 1-3 [5]. fayez tarsha kurdi (2021), random forest machine learning technique for automatic vegetation detection and modelling in lidar data. international journal of environmental sciences & natural resources. pp. 001. (fayez tarsha [6]. juan a. herrera-silva and myriam hernández-álvarez (2023), dynamic feature dataset for ransomware detection using machine learning algorithms. pp. 1-21 [7]. sh kok, azween abdullah, nz jhanjhi and mahadevan supramaniam (2019) ransomware, threat and detection techniques: a review. ilcsns international journal of computer science and network security, vol. 19.2, pp. 138-139. [8]. kok s.h. and mahadevan (2019) prevention of crypto-ransomware using a pre-encryption detection algorithm. articles www.mdpi.com/journal/computers. pp. 2-5 [9]. manabu hirano and ryotaro kobayashi (2019) ‘machine learning based ransomware detection using storage access patterns obtained from live-forensic hypervisor” conference paper · october 2019. pp. 2-7. [10]. olaniyi abiodun ayeni, otasowie owolafe, olabiyi akinsola (2021), malware detection using machine learning, conference paper. pp. 86 [11]. samah alsoghyer and iman almomani (2019) ransomware detection system for android applications. pp. 1-31. [12]. sh kok, azween abdullah, nz jhanjhi and mahadevan supramaniam (2019) ransomware, threat and detection techniques: a review. ilcsns international journal of computer science and network security, vol. 19.2, pp. 138-139. [13]. sh kok, azween abdullah, nz jhanjhi and mahadeyan supramaniam. (2019), ransomware, threat and detection techniques: a review. pp. 1-11 [14]. subash poudyal, dipankar dasgupta, zahid akhtar, kishor datta gupta., (2019), a multi-level ransomware detection framework using natural language processing and machine learning. pp. 2-9 [15]. xiang g., et al. (2019). an improved random forest algorithm for predicting employee turnover. research article. pp. 2-5 http://oreilly.com/catalog/errata.csp?isbn=9781449303716 http://www.mdpi.com/journal/computers.%20pp.%202-5 vol. 6, no.1, january 2025 | 18 p-issn: 2715-2448 | e-issn: 2715-7199 vol.4 no.1 january 2023 buana information technology and computer sciences (bit and cs) design of a vehicle parking management information system in tanjung duren central park area west jakarta topan setiawan komputerisasi akuntansi, universitas ma’soem, indonesia e-mail: topansetiawan@masoemuniversity.ac.id received: 2024-06-01 | revised: 2024-12-10 | accepted: 2025-01-11 abstract tanjung duren parking is a parking lot whose operational activities still use a manual system, so that parking officers often experience problems such as incorrectly calculating the total parking fee or incorrectly recording incoming and outgoing vehicles, which results in parking revenue anomalies and inaccurate parking space availability data. therefore, in this research, a parking vehicle management information system was designed as a solution to existing problems. this information system was designed using the system development life cycle (sdlc) method and approach. based on the results of research that has been carried out, it is known that the information system that has been designed is effective to implement because it is able to solve existing problems such as automatically calculating parking rates and recording vehicles both when entering and leaving. keywords: information system design, parking management information system, tanjung duren parking i. introduction parking is defined as the process of placing a vehicle with the aim of stopping or stopping the vehicle temporarily [1]. parking is usually carried out by vehicle drivers in places such as public parking lots, parking lots in shopping centers, office buildings, or on the side of the road, where the purpose of parking can vary, from meeting temporary needs to being part of a long-term travel plan [2]. with the increasing number of motorized vehicle users, especially in urban areas, parking spaces have become an important element, this is because apart from being part of a safe and organized urban infrastructure, parking spaces can also help reduce road congestion, improve vehicle safety, and facilitate activities in urban centers [3]. apart from that, parking lots can also support the use of public transportation and optimize the use of city space [4]. parking lots are typically managed by various entities, including local governments, private companies, or independent parking operators. local governments often have a role in regulating and managing parking in public spaces, such as on the side of the road or in public areas. they are responsible for establishing parking policies, setting parking rates, and ensuring safety and order in parking management. on the other hand, private companies or independent parking operators can manage parking in shopping centers, offices, or entertainment centers by providing parking services, maintenance, and security monitoring. in some cases, partnerships between local governments and the private sector can also be formed to manage parking more efficiently, such as the parking partnership carried out by an independent parking operator in the tanjung duren central park area, west jakarta, which is named tanjung duren parking [5]. tanjung duren parking is a special parking area for motocycle, where in its operational activities the parking lot in this area can accommodate more than 500 motocycles with a plate fee of idr 5,000 per day. currently, the system used still uses a manual system where vehicles wishing to park will be given an entry parking ticket, then payment will be made when the vehicle is about to leave. currently, the parking management system used is still conventionally based, where the calculation of the availability of empty parking areas is based on the number of tickets sold, incoming vehicles and also vol. 6, no.1, january 2025 | 19 outgoing vehicles [6]. when the intensity of vehicle entry and exit is quite high, parking officers often encounter problems such as incorrectly calculating the total parking fee or incorrectly recording the entry and exit of vehicles. the direct impact of this error, apart from causing anomalies in parking revenue, also causes parking space availability data to be inaccurate. as a result, it often takes drivers longer to look for an empty parking area, and this of course has an impact on long queues of vehicles that spill onto the main road, causing traffic jams [7]. therefore, we need a parking management information system that can provide fast, precise and accurate information, and can be a solution to current problems. ii. methods method is a systematic method used by a researcher to solve or find answers to the problems being faced in research. the method used in this research refers to the stages of the system development life cycle (sdlc) method. in information systems engineering, sdlc is the process of creating and modifying systems as well as the models and methodologies used to develop these systems. this model is also a reference used as a framework in research [8], [9]. figure 1. research concept framework based on figure 1, it can be seen that there are five stages that form the research framework. the details of each step are as follows: 1. data analysis and collection at this stage, all necessary information is collected to understand the needs and requirements of the system to be built. this activity includes interviews with stakeholders, surveys, observations, and analysis of existing documents. 2. data processing and requirements estimation the data collected in the previous stage is then further analyzed to identify the technical and functional requirements of the system. requirement estimation involves determining the necessary resources, such as time, cost, manpower, and the technology to be used. 3. system modeling and design based on the specified requirements, this stage focuses on modeling and designing the system. this includes creating flow diagrams, entity-relationship diagrams, user interface designs, and system architecture designs. vol. 6, no.1, january 2025 | 20 4. system implementation after the design phase is completed, the next step is to build the system or application according to the specified design. 5. report and publication this stage is where the system is implemented. it also involves documenting the project, including test results, user manuals, and the final project report. iii. results and discussions 1. current system analysis analysis is carried out with the aim of observing an object by breaking it down into smaller parts, looking for weaknesses and then proposing improvements. based on the research that has been carried out, the current parking management system is as follows: 1. upon entry into the parking area, the parking attendant will record the vehicle's license plate number and entry date on the parking ticket and then hand it to the driver. 2. the attendant will then log the vehicle's entry into the daily record book. 3. upon exit, the parking attendant will check the driver's parking ticket, verify the entry date, and calculate the parking fee to be paid. 4. the driver will then pay the parking fee as charged. 5. finally, the attendant will log the vehicle's exit into the daily record book. figure 2. current system flowmap based on the flowmap above, several problems can be identified as follows: 1. staff need time to calculate parking rates if the vehicle is parked for more than one day. 2. staff have the potential to forget to check incoming or outgoing vehicles. 3. daily log book or parking ticket are vulnerable to damage or loss. vol. 6, no.1, january 2025 | 21 2. proposed work system based on the analysis that has been carried out on the current system, and identification of existing problems, the proposed system has the following working procedures: 1. the attendant enters the vehicle's number into the system and saves it. 2. afterward, the system automatically prints a parking ticket with the entry date and time, complete with a barcode, and hands it to the driver. 3. when the vehicle is ready to exit, the driver provides the parking ticket to the attendant for scanning, and the system calculates the parking fee based on the vehicle's entry date and time. 4. the driver then pays the parking fee as charged. 5. the attendant saves the vehicle's exit data while printing a payment receipt and hands it to the driver. 6. reports can be generated daily, weekly, or monthly in two copies: one copy is given to the owner, and the second copy is kept as an archive. figure 3. proposed work system flowmap 3. context diagram context diagram is a high-level, simplified visual representation of a system and its interactions with external entities. it provides an overview of how the system interfaces with outside actors like users, other systems, or external organizations. this diagram captures the system as a single process, highlighting its boundaries and the data flow between the system and external entities. the primary purpose is to establish the context and scope of the system, ensuring all external interactions are clearly identified [10]. based on the flowmap in figure 2, it can be seen that this parking vol. 6, no.1, january 2025 | 22 management information system has two external entities, namely pengendara (driver) and pemilik (owner), and one internal entity, namely petugas (officer). figure 4. context diagram 4. data flow diagram data flow diagram (dfd) is a graphical representation used to depict the flow of data within a system. it illustrates how data moves through the system, showing inputs, processes, data stores, and outputs [11]. the dfd breaks down the system into smaller components to display how data is processed and transferred at various stages. it helps in understanding the functional aspects of the system and how different components interact through the flow of data. dfds are typically created at multiple levels, with high-level dfds providing an overview and detailed dfds depicting specific processes [12]. the dfd of this parking management information system looks like in figure 4. figure 5. data flow diagram 5. entity relationship diagram entity-relationship diagram (erd) is a structural diagram used to model the data relationships within a database [13]. it visually represents entities (objects or concepts) and the relationships between them. relationships define how entities interact with each other. figure 6. entity relationship diagram 6. program menu structure program menu structure is a hierarchical layout that organizes the menus and navigation paths within a software application. it outlines how menus and sub-menus are arranged, showing the structure of the user interface [14]. the main menu provides access to the primary functions or sections of the application, while sub-menus offer more specific options or features. each menu item performs a particular action or command. the menu structure is designed to be intuitive, allowing users to navigate the software easily and access all necessary functions efficiently. this structure is sistem informasi pengelolaan parkir pengendara pemilik tiket parkir dan uang tunai tiket parkir dan resi bay ar laporan parkir pengendara 1.0 parkir masuk 2.0 parkir keluar 3.0 pembuatan laporan tiket parkir tiketparkir data kendaraan masuk ambil data data kendaraan keluar tiket parkir resi pembay aran ambil data pemilik laporan parkir petugas nik** (char5) namapetugas (varchar25) alamat (varchar50) telepon (varchar13) petugas noparkir** (varchar20) nopolisi (varchar10) tanggaljammasuk (datetime) tanggaljamkeluar (datetime) tarif parkir (integer) nik* (char5) tikerparkir 1 n vol. 6, no.1, january 2025 | 23 crucial for creating a user-friendly interface and improving overall usability. the application program menu structure of this parking management information system looks like in figure 6. figure 7. program menu structure 7. form password this form will be displayed for the first time. petugas is required to enter their nik and password when logging into the application. this form is used as protection between one officer and another officer. figure 8. form password 8. form parkir masuk dan keluar kendaraan when the vehicle is about to enter, petugas can enter the vehicle plate number or police number into the text box provided, then save it. the system will then print a tiket parkir according to the date and time of entry which is equipped with a barcode. meanwhile, when leaving, petugas can enter the no. parkir by inputting it manually or using a scanner, then the program will calculate the parking rate that must be paid by pengendara. data mastersistem transaksi laporan keluar menu utama log in log out data petugas kendaraan masuk kendaraan keluar vol. 6, no.1, january 2025 | 24 figure 9. form parkir masuk dan keluar dan tiket parkir 9. parking report the parking management report is an output that is processed by the application program and contains information that can be used as a basis for decisions both now and in the future [15]. figure 10. parking report iv. conclution based on the research that has been carried out, it can be concluded that this vehicle parking management information system can be implemented as a solution to existing problems, such as being able to automatically calculate parking rates and being able to record vehicles both when entering and leaving. there are suggestions for future researchers so that the research can be developed by adding other features such as recording cases of loss, damage, and others. references [1] t. a. insani, “perancangan sistem informasi pengelolaan parkir pada pt. disire jaya mandiri berbasis web menggunakan metode personal extreme programming (pxp),” teknol. inf. esit, vol. xvii, no. 01, pp. 48–53, 2022. [2] a. hidayat and a. maskhun, “sistem informasi parkir kendaraan berbasis android di pt piranti indonesia,” j. manaj. inform., vol. 8, no. 2, pp. 43–52, 2021, doi: 10.51530/jumika.v8i2.557. [3] a. b. warsito, m. yusup, and m. aspuri, “penerapan sistem monitoring parkir kendaraan berbasis android pada perguruan tinggi raharja,” technomedia j., vol. 2, no. 1, pp. 82–94, 2017, doi: 10.33050/tmj.v2i1.317. [4] amuharnis and a. rahmadian, “aplikasi pengelolaan parkir kendaraan dengan menentukan blok parkir,” j. coreit j. has. penelit. ilmu komput. dan teknol. inf., vol. 3, no. 1, pp. 35–40, 2017, doi: 10.24014/coreit.v3i1.3717. vol. 6, no.1, january 2025 | 25 [5] z. millenia and a. farida, “tata kelola pendapatan parkir: petugas parkir dishub dan non dishub,” j. manaj. dan penelit. akunt., vol. 15, no. 2, pp. 105–120, 2022, doi: 10.58431/jumpa.v15i2.202. [6] m. taufiq ramadhan, d. rizki rahadian, and dkk., “perancangan aplikasi parkir kendaraan berbasis website dengan metode waterfall,” maklumatika, vol. 10, no. 1, pp. 10–19, 2023. [7] m. n. k. nababan, t. d. s. rumapea, and dkk., “pemodelan sistem parkir kendaraan berbasis android menggunakan algoritma aes,” jikomsi (jurnal ilmu komput. dan sist. informasi), vol. 3, no. 2, pp. 76–80, 2020, [online]. [8] a. neno, a. yuniar rahman, and f. marisa, “a detection of malacca woven fabric motifs using the yolov4 method,” buana inf. technol. comput. sci. (bit cs), vol. 5, no. 1, pp. 45– 50, 2024, doi: 10.36805/bit-cs.v5i1.6081. [9] m. nisti, a. yuniar rahman, and f. marisa, “detection of diseases and pests on the leaves of sweet potato plants sing yolov4,” buana inf. technol. comput. sci. (bit cs), vol. 5, no. 1, pp. 39–44, 2024, doi: 10.36805/bit-cs.v5i1.6065. [10] h. yunita and dina, “aplikasi pelayanan kesehatan pada puskesmas,” jitek (jurnal inform. dan teknol. komputer), vol. 1, no. 1, pp. 1–13, 2021. [11] m. irfan, d. mirwansyah, and dkk., “perancangan sistem informasi monitoring akademik dengan menggunakan data flow diagram,” j. locus penelit. dan pengabdi., vol. 2, no. 12, pp. 1201–1207, 2023, doi: 10.58344/locus.v2i12.2352. [12] m. a. triyandi and muntahanah, “implementasi location based service dalam perancangan aplikasi pencarian lokasi toko grosir manisan di kota bengkulu,” j. media infotama, vol. 20, no. 1, pp. 264–273, 2024. [13] k. ’afiifah, z. f. azzahra, and dkk., “analisis teknik entity-relationship diagram dalam perancangan database sebuah literature review,” inform. dan teknol., vol. 3, no. 1, pp. 08– 11, 2022, doi: 10.54895/intech.v3i2.1682. [14] t. setiawan, a. suryopratomo, and dkk., “rancang bangun sistem informasi media pembelajaran interaktif bahasa inggris berbasis web,” j. account. inf. syst., vol. 5, no. 2, pp. 186–196, 2022, doi: 10.32627/aims.v5i2.554. [15] j. h. p. sitorus and m. sakban, “perancangan sistem informasi penjualan berbasis web pada toko mandiri 88 pematangsiantar,” j. bisantara inform., vol. 5, no. 2, pp. 1–13, 2021. vol. 6, no.1, january 2025 | 39 implementation of the c4.5 algorithm in predicting the interest of prospective students in choosing higher education bambang triraharjo1*, prilian ayu minarni2, baskoro3 1,3universitas muhammadiyah pringsewu/fakultas ilmu komputer, pringsewu,lampung, indonesia 2universitas muhammadiyah pringsewu/fakultas kesehatan, pringsewu,lampung, indonesia e-mail: bambangtriraharjo@umpri.ac.id, prilianayuminarni@umpri.ac.id, baskoro@umpri.ac.id received: 2024-07-02 | revised: 2024-10-30 | accepted: 2025-01-31 abstract lampung province has many diverse private universities and offers a variety of majors. due to the large number of private universities that exist, competition between universities to attract prospective new students is very tight. so, to be able to compete with these universities, the campus needs to predict the interests of prospective new students by knowing what factors motivate prospective new students to choose a university. the aim of this research is to predict prospective students' interest in choosing a university. in this research, the data processed are the results of a survey conducted on prospective new students in the information systems and technology undergraduate study program. data from prospective students will be processed using a data mining process with the classification method. furthermore, using the c4.5 algorithm in the classification method, literacy was obtained up to node 4 with 3 assessment data which became a factor in determining prospective students' interest in the study program. furthermore, the results of data processing using the classification method with the c4.5 algorithm were tested using rapid software. miner, where an accuracy rate of 100% is obtained. the final result of using data mining with the c4.5 algorithm is that it is able to predict the interests of prospective new students based on assessment factors in choosing a university. keywords: prediction, classification, data mining, c4.5 algorithm, decision tree i. introduction the development of a university, one of which is seen based on the number of students obtained every year. in universities, especially private, the acquisition of students every year is the main factor in the development of the university. because, students are the main resource for private universities to get funds in carrying out operations in universities. the more students, the greater the income for the university, so that it is easy for the university to run campus operations without limited funds. currently, the information systems and technology study program, university of muhammadiyah pringsewu is a new program majoring in computer science. in 2021-2022, the information systems and technology study program will only start opening registration. so currently the study program is trying its best to get prospective new students to be able to continue their education at the information systems and technology study program, university of muhammadiyah pringsewu. in choosing a university, a student will usually look for information about the university they are going to [8]. apart from that, there are also factors that affect prospective students in choosing a university, including factors that influence parents, friends, relatives, scholarships, job opportunities and others. with these factors, universities can predict what factors are the main drivers for prospective students to choose universities, so that universities can make the best decisions to be able to recruit as many prospective students as possible [3]. interest can be interpreted as a high tendency and passion or a great desire for something. the purpose of this study is to find out the interest of prospective new p-issn: 2715-2448 | e-issn: 2715-7199 vol.6 no.1 january 2025 buana information technology and computer sciences (bit and cs) mailto:prilianayuminarni@umpri.ac.id vol. 6, no.1, january 2025 | 40 students in a university and how to apply the data mining process to predict the request of prospective students based on the assessment factors they use [6]. data mining can also be interpreted as the extraction of new information extracted from chunks of big data that helps in decision-making [15]. data mining is part of knowledge discovery in database (kdd) which consists of several stages such as data selection, data pre-processing, transformation, data mining and evaluation of results [11]. by using data mining, valuable information in the data set can be mined. by using data mining, data on the interests of prospective new students can be processed with an algorithm [10]. in data mining there are algorithms that are able to analyze data. in this study, the algorithm used is the c4.5 algorithm. the c4.5 algorithm is a classification and prediction method used to form a decission tree based on training data [9]. a decission tree is a flowchart structure that has a tree, where each internal node indicates an attribute test, each branch represents the test result and the leaf node represents the class or class distribution [7]. in the previous study, the theory of planned behavior approach was used in predicting student interest[5]. this approach measures students' interest in entrepreneurship by prioritizing 3 determinants of desire to be entrepreneurial, namely attitudes towards behavior, subjective norms, and behavior control. the result of this approach is a conclusion where attitudes towards behavior do not affect students' interest while 2 other factors do [6]. another algorithm used in previous research to predict the interest of vocational school students to enter college is the naive bayes algorithm. the data used is graduate data in 2018 and 2019 where there are 158 data. furthermore, the data was processed using the naive bayes algorithm with 9 attributes, namely expertise, report card scores, un scores, parental work, and others. the data was tested using the rapidminner application [2]. the results of the processor were obtained from the prediction of parental income attributes, report card scores, and student desires, which affected students' interest in continuing their education to higher education with an accuracy rate of 92.96% [12]. in previous research, the c4.5 algorithm has been used to predict the factors that cause students to repeat a course. in the study, variables were determined in the form of number of semesters, gpa, grades, economic condition and status. furthermore, data processing is carried out with the c4.5 algorithm as many as 141 data records as training data. after data processing using the c4.5 algorithm, tests were then carried out using the weka application. the result of this study is information in the form of rules that can predict students who repeat a course [1]. furthermore, the c4.5 algorithm is also used for the classification of student success rate predictions at amik tunas bangsa. in this study, several attribute variables were determined such as gender, attendance, lecture session, average score and school origin. furthermore, data processing was carried out using the c4.5 algorithm. after data processing using the c4.5 algorithm, the results were tested using rapid minner. the results of the processing of the rapidminner application found that the results of decisions obtained based on manual calculations and using rapid minner resulted in an accuracy of 92% against predictions [14]. from the previous research that has been explained, the use of the theory of planned behavior approach has not been optimal in helping decision-makers. this is because the prediction accuracy level is not accurate, while the naive bayes algorithm compared to the c4.5 algorithm has a low accuracy level, and also does not describe the shape of a decision with a decession tree. with the use of the c4.5 algorithm, the data will be tested with a high degree of accuracy to predict the interest of prospective new students in choosing a university so that the university can make the best decision in attracting prospective new students. ii. methods in this study, the stages in the process of getting the best decision using the c4.5 algorithm based on the data that have been obtained are explained. the stages carried out can be seen in figure 1. vol. 6, no.1, january 2025 | 41 figure 1. research stages a. data needs analysis the data needed in this study is data in the form of survey results data on students and criteria data that are factors in determining prospective students interested in studying at the campus. there are 5 criteria and subcriteria that are the determining factors that can be seen in table 1. table 1. criteria for determining factors of interest in prospective students choosing a campus it criterion subkriteria 1 tuition cheap affordable expensive 2 facilities excellent adequate enough 3 accreditation excellent good enough 4 service excellent good enough 5 location near affordable far the reason for choosing the criteria for accreditation, services and facilities is seen in terms of student needs for higher education. in terms of tuition fees, because many families in lampung have a lower middle monthly income and in terms of location, many students come from outside the lampung area. based on the criteria of the above determining factors, a survey was conducted on several students based on the criteria of these determining factors so that data on the interest survey of students who chose and did not choose to study at the campus were obtained. furthermore, the data from the survey results was carried out data transformation for processing. b. data transformation data transformation is a change in data by changing the initial attribute value into an attribute value that is in accordance with the needs of the data in its processing [4]. based on the results of the survey of prospective students' interest in choosing a campus against the criteria, then data transformation was carried out. for testing using the c4.5 algorithm, 25 sample data is used as a training dataset. the results of the survey data transformation can be seen in table 2: table 2. data transformation based on survey data mhs accreditat ion tuition facilities service location interest mhs 1 good affordable excellent good affordable yes mhs 2 enough cheap enough enough affordable yes mhs 3 enough affordable enough enough far it vol. 6, no.1, january 2025 | 42 mhs 4 excellent cheap excellent excellent affordable yes mhs 5 good cheap adequate good far yes mhs 6 excellent affordable adequate good affordable yes mhs 7 enough affordable enough enough far it mhs 8 good affordable enough good near yes mhs 9 enough expensive enough enough affordable it mhs 10 enough affordable enough enough far it mhs 11 good cheap excellent excellent near yes mhs 12 enough affordable excellent enough far yes mhs 13 enough affordable adequate good affordable yes mhs 14 enough expensive adequate good near it the data contained in the data transformation is sample data that will be used for testing using the c4.5 algorithm. c. c4.5 algorithm process based on the results of the data transformation that has been obtained, the c4.5 algorithm process is then carried out. the c4.5 algorithm has stages or steps in the problem-solving process so that it reaches the stage of results. the stages in the c4.5 algorithm process can be seen as follows [13]: 1. calculate the number of "yes" and "no" cases based on transformation data. 2. determine the entropy value of each criterion on a case-by-case basis. to calculate the entropy value can use the formula: entropy (s) = ∑𝑛 i=1 − 𝑝𝑖 ∗ 𝑙𝑜𝑔2𝑝𝑖 (1) where: s = case set n = number of partitions s pi = proportion of si to s 3. determine the highest gain value based on the entropy score that has been obtained. to calculate the gain value you can use the formula: gain(s,a) = entropy(s) – ∑𝑛 i=1 |si/s| * entropy(si) (2) a. where: b. s = case set c. a = features d. n= number of partitions attribute a |si| = the proportion of si to s |s|= number of cases in s 4. generate a decission tree to generate the first node, go through processes 1 to 4 until there are no records in the empty decission tree branch. iii. results and discussions a. data processing based on the algorithm process described above, then testing is carried out using data mining software, namely rapidminer. for testing using the rapidminer application, 592 data testing data was used. testing is carried out by importing test data in excel format into rapid miner software as shown in figure 2. vol. 6, no.1, january 2025 | 43 figure 2. dataset snippets decision tree based algorithm on the imported data, drag and drop the learning dataset table into the process view and set the decession tree operator which can be seen in figure 3: figure 3. process view in rapidminer software based on the process view that has been made, the acquisition of entropy and gain values from node 4, then to obtain the gain value of the entire category is 0, then for the service and location categories are not a factor in determining student interest in choosing a campus. the shape of the final decession tree can be seen in figure 4: figure 4. decession tree test results from the results of the test using rapidminer using the decession tree algorithm, it is explained that the criteria of facility simply have a "no" decision for students' interest. so the decession tree above is used as the final decision in the prediction of determining the request for prospective students in choosing a university. vol. 6, no.1, january 2025 | 44 b. evaluate results the results of the decession tree from the c4.5 algorithm are then evaluated on the results of the decession tree based on the transformation data. the evaluation aims to see if there are errors in the results of the decession tree obtained and whether it is necessary to re-test. it is then translated into the form of decisions. based on the translation of the decision, it can be a tool for a decision-maker to predict what criteria factors support the interest of prospective students in choosing the desired university. so that the prediction can be taken by the campus to take appropriate actions in dealing with this problem. the form of the decision generated through the rapid miner software can be seen in figure 5: figure 5. test results based on the tests that have been carried out, it can be concluded that the main factors that affect the interest of prospective new students in choosing a university using the c4.5 algorithm are as follows: 1. if the tuition fee is "cheap", then the interest of prospective students is "yes" to choose a university and if the tuition fee is "expensive", then the interest of prospective students is "no" to choose a university. 2. if the tuition fee is "affordable" and the campus accreditation is "very good" or "good", then the interest of prospective students is "yes" to choose a college. 3. if the tuition fee is "affordable" and the campus accreditation is "sufficient" and the campus facilities are "very good" or "adequate", then the interest of prospective students is "yes" to choose a university, but if the tuition fee is "affordable" and the campus accreditation is "sufficient" and the campus facilities are "adequate", then the interest of prospective students is "no" to choose a university. iv. conclusions based on the results of the research obtained, it is concluded that the data mining process using the c4.5 algorithm has been able to produce a decision to predict the interest of prospective new students in choosing universities based on the criteria factors. based on the results of data processing using the c4.5 algorithm with 25 sample data, the criteria that are the main factors for the interest of prospective new students in choosing a university are the tuition, accreditation and facilities criteria, where cheap tuition, excellent or good accreditation and very good or adequate facilities attract the interest of prospective new students in choosing a university. meanwhile, expensive tuition, sufficient accreditation and sufficient facilities do not attract the interest of prospective new students in choosing a university. based on testing using rapid miner software with 592 data, the results of the decession tree and the decision between manual testing are the same, with an accuracy level of 100%. so that the use of the c4.5 algorithm has produced an assessment factor that can predict the interest of prospective students, and the university can predict the interest of prospective new students in the future. vol. 6, no.1, january 2025 | 45 references [1] alwarthan, s. a., aslam, n., & khan, i. u. (2022). predicting student academic performance at higher education using data mining: a systematic review. in applied computational intelligence and soft computing (vol. 2022). hindawi limited. [2] berrar, d. (2019). bayes’ theorem and naive bayes classifier. in s. ranganathan, m. gribskov, k. nakai, & c. schönbach (eds.), encyclopedia of bioinformatics and computational biology (pp. 403–412). academic press. [3] feng, l. (2021). research on higher education evaluation and decision-making based on data mining. scientific programming, 2021. https://doi.org/10.1155/2021/6195067 [4] galit shmueli, p. c. b. i. y. n. r. p. k. c. l. jr. (2018). data mining for business analytics. [5] gerhana, y. a., fallah, i., zulfikar, w. b., maylawati, d. s., & ramdhani, m. a. (2019). comparison of naive bayes classifier and c4.5 algorithms in predicting student study period. journal of physics: conference series, 1280(2). [6] hu, j., & li, h. (2021). composition and optimization of higher education management system based on data mining technology. scientific programming, 2021. [7] masters, t. (2018). data mining algorithms in c++. in data mining algorithms in c++. apress. https://doi.org/10.1007/978-1-4842-3315-3 [8] naga, j. f., & tinam-isan, m. a. c. (2024). exploring the influence of personality traits on students’ information security risk-taking behaviors: a bfi assessment. procedia computer science, 234, 527–536. https://doi.org/10.1016/j.procs.2024.03.036 [9] ngoc, p. v., ngoc, c. v. t., ngoc, t. v. t., & duy, d. n. (2019). a c4.5 algorithm for english emotional classification. evolving systems, 10(3), 425–451. [10] nurmalitasari, awang long, z., & faizuddin mohd noor, m. (2023). factors influencing dropout students in higher education. education research international, 2023. [11] parteek bhatia. (2019). data mining and data warehousing. [12] wang, c. (2021). analysis of students’ behavior in english online education based on data mining. mobile information systems, 2021. https://doi.org/10.1155/2021/1856690 [13] wang, j. (2022). application of c4.5 decision tree algorithm for evaluating the college music education. mobile information systems, 2022. https://doi.org/10.1155/2022/7442352 [14] yang, x., & ge, j. (2022). predicting student learning effectiveness in higher education based on big data analysis. mobile information systems, 2022. https://doi.org/10.1155/2022/8409780 [15] ye, h., & li, c. (2022). engineering education understanding expert decision system research and application. computational intelligence and neuroscience, 2022. vol.6, no.2, july 2025 | 100 p-issn: 2715-2448 | e-issn: 2715-7199 vol.6 no.2 july 2025 buana information technology and computer sciences (bit and cs) e-konsulta: clinic medical record management system with online follow-up through google meet rianna marie c. batersal1*, omar bin ayob d. cadingilan2, roselyn a. gulbe3 philipcris c. encarnacion4, jonel e. ebol5 1,2,3,4,5 college of computing studies saint columban college pagadian city, philippines e-mail: riannamarie.batersal@sccpag.edu.ph1, omarbinayob.cadingilan@sccpag.edu.p2, roselyn.gulbe@sccpag.edu.ph3, philcrisen@sccpag.edu.ph5, jonel.ebol@sccpag.edu.ph6 received: 2025/02/02 | revised: 2025/07/02 | accepted: 2025/07/29 abstract an operational clinic is essential in educational institutions to support the health and well-being of students and staff, directly enhancing productivity and academic achievement. however, many clinics still rely on outdated, paper-based systems, leading to inefficiencies, delays, and errors. modernizing healthcare through digital solutions can improve efficiency, accuracy, and accessibility. this study proposed a system application to address common issues by digitizing medical records, automating workflows, and enabling online consultations. the platform enhances the accuracy and accessibility of patient data while streamlining administrative tasks, benefiting all stakeholders. extensive testing revealed excellent performance, with a 100% pass rate in most functional areas and high scores in nonfunctional aspects, including reliability (97.14%), security (95.71%), and user experience (98.57%). minor areas for improvement were noted in certain functionalities (92.90%) and accessibility (90%), underscoring the importance of continuous optimization to enhance inclusivity and usability. future enhancements could include developing a mobile application for improved accessibility, enabling users to access records, book appointments, and participate in consultations on their devices. advanced features such as enhanced online follow-up consultations, automated reminders, and health monitoring tools could further improve healthcare delivery. integrating real-time analytics or telehealth capabilities may also provide broader support for patient care. these recommendations highlight the potential for the system to evolve, offering a more efficient, inclusive, and responsive approach to healthcare within educational environments. keywords: health management system, medical record, online follow-up, google meet. i. introduction an operational clinic is essential in educational institutions to support the health and well-being of students and staff, directly enhancing productivity and academic achievement. however, many clinics still rely on outdated, paper-based systems, leading to inefficiencies, delays, and errors. modernizing healthcare through digital solutions can improve efficiency, accuracy, and accessibility. a digital healthcare platform addresses these challenges by streamlining workflows, managing patient records, tracking and inventory, and ensuring accurate, confidential, and timely healthcare services. research shows that international perspectives highlight the importance of accessible healthcare in schools, with studies indicating that early detection of health issues significantly enhances overall well-being and academic success [1]. research has also shown that adopting digital solutions in healthcare can improve service delivery, reduce administrative burdens, and ensure more effective health interventions [2]. additionally, innovative tools such as telehealth and automated health systems have been associated with increased patient satisfaction and operational efficiency [3]. this project mailto:riannamarie.batersal@sccpag.edu.ph1 mailto:omarbinayob.cadingilan@sccpag.edu.p2 mailto:roselyn.gulbe@sccpag.edu.ph3 mailto:philcrisen@sccpag.edu.ph5 mailto:jonel.ebol@sccpag.edu.ph vol.6, no.2, july 2025 | 101 reflects a commitment to leveraging advanced technologies to overcome logistical barriers, enhance patient care, and foster a healthier, more productive educational environment. an international study emphasizes the importance of accessible healthcare within schools to improve outcomes and enhance the capabilities of both students and staff. this perspective is supported by the world health organization (who), which recognizes that such an approach is highly effective in promoting the overall wellbeing of students in the school environment [4]. the project we developed turned out to be an invaluable tool for addressing the challenges the school clinic staff faced. prior to implementation, managing patient records, appointments, and medical supplies often created significant inefficiencies and stress. the study is sourced by a diverse range of sources, ultimately leading to developing a system that offers a novel and innovative approach for our clients to manage their resources and patients more effectively. the design of this system is grounded in the principles of electronic health records (ehrs), which are recognized for their ability to reduce documentation errors, enhance decision-making, and facilitate efficient health data management [5]. addressing this systematic problem is truly innovative, essential, and beneficial for equitable access to healthcare for students and employees [2]. thus, by adapting an approach like this it could be a mean to overcome logistics challenges in which everyone in the school premises can benefit into this high-quality healthcare service. as shown in the figure 1 below is our product perspective in which it illustrates and highlights the seamless integration of the business logic of the school clinic which rendered our approach in addressing the institutes challenges. figure 1. product perspective to further expand our vision for this study, our project is not just aiming for the proper organization of the seamless, paperless transaction of the clinic staff; we, as a team, developed a system in which online consultation and followup can be reached with the help of google meet. as shown in the figure 1 the transaction between the system and the google meet is made through by the help of an api in which it allows us to generate the code and link and redirect the clinic user by this case e.g. nurse or dentist to the browser to finally have a proper setup of their scheduled online follow-up consultation. vol.6, no.2, july 2025 | 102 ii. methods the research methods employed in developing our project were based on the waterfall development model [6] [7]. this model provided a structured and sequential approach, allowing us to define and complete each phase of the development process systematically. additionally, the project was developed using the c# programming language, which enabled us to work in a familiar and efficient environment. while the foundation of the study is the waterfall model, we incorporated elements of both iterative and incremental approaches. this combination allowed us to adapt to client requirements more effectively and refine the system progressively. by leveraging these approaches, we ensured the delivery of a high quality and efficient software product that meets the client’s needs. to further support our study, we used a descriptive approach on our research method in which it enables us to utilize our survey questionnaire as a tool in gathering important data to further access the improvement, updates and functionalities that is needed to be implemented on our system. an addition to that, according to a study and the reason we implemented a descriptive method on our study its because it allows us researchers to describe the characteristic of certain population to be studied further and be able to narrow our scope for a precise delivery of the satisfaction towards our own client [8]. the development of the said project consist of the following stages that are based on the waterfall model: 1. requirements gathering during this phase, we gathered important data, such as the clinic's business logic in which we come up with the final product perspective on the figure 1 and necessary data specifically related to clients' needs. identifying the project’s objective and requirements is very crucial to ensure that the clarity and alignment of the system are aligned with the client's expectations. 2. system design right after gathering the requirements, we, the researchers, created a detailed design for the system, in which we were able to come up with our studied use case diagram below on the figure 2 it illustrates the whole concept of the system, breaking down and aligning with the business logic of our clients' needs. we also implemented the google meet code for online follow-up consultations. figure 2: use case diagram vol.6, no.2, july 2025 | 103 3. implementation the development of the system is developed using the c# programming language, which also follows the concept of our client and follows the standard object-oriented program and umls to have the standard result [7],[10],[11]. we researchers use other tools such as figma, canva, and visual studio as an environment for developing the system and enhancing the system's beautification and functionality. the development also included mysql as our database. since our system mainly focuses on offline and local data storage, it is recommended that we use mysql since it is more familiar to us researchers and will be able to manage the system's creation. each module of the system was coded according to the design specifications, which can be referred to as the figure 3 below. figure 3: module interface (dashboard) 4. testing at this juncture, we researchers proceed with the testing of the system with rigorous methods such as aligning the test survey and questionnaire with the iso standard 25010 and being able to gather information for the improvements of both the functional and non-functional requirements of the system [9],12]. the study also included unit testing, which helps us researchers to identify the bugs and lines of code that need to be refactored. since we are using c#, we conducted the unit test is based on nunit testing, in which we only identify the important modules that are most likely to be used throughout the system. below on the error! reference source not found. these are the modules of our system that have been tested. we likely chose these modules to be tested because nunit is only available with methods. since our system is more likely focused on the ui and databases, we chose these modules because they only have the important methods that need to be used. table 1: nunit testing final result category test name duration dashboardtest expired_whencalled_datag ridview4ispopulatedwithex pecteddata 3 second sched_whenexception_mess ageboxshowsinvalidlogin 19 millisecond viewss_whenexception_me ssageboxshowsinvalidlogin 6 millisecond views_whenexception_mess ageboxshowsinvalidlogin 5 millisecond xtodayspatients_whenexcept ion_messageboxshowsinval idlogin 6 millisecond vol.6, no.2, july 2025 | 104 xtxodaysschedule_whencall ed_setslabeltexttocount 4 millisecond xtxodaysschedule_whenexc eption_messageboxshowsin validlogin 4 second xxtodayspatients_whenexce ption_messageboxshowsinv alidlogin 16 millisecond xxweekpatients_whenexcept ion_messageboxshowsinval idlogin 10 millisecond form1tests auto6_shouldpopulateauto completecollection_whend atabasehasvalues 3 second button4_click_shouldnotsh owmessage_whenallfields arefilled 899 millisecond button4_click_shouldshow message_whenfieldsareem pty 762 millisecond button4_click_shouldshow message_whenonefieldisfi lled 6 second logintest login_shoulddisplaysucces smessage_whenvalidcreden tialsareentered 47 millisecond the table above shows the total value of how fast each module will be able to respond to each test that it is required to. this information helps researchers decide how to refactor the codes and functionality that need to be improved or changed [13]. 5. deployment right after the testing method, we proceeded with the deployment of the system. the system was deployed successfully with the client's permission, and now the system is in its working phase. this phase involves setting up the proper environment of the system and the connectivity of the system on each personal computer in which it involves networking via wireless connectivity. we researchers ensure that the system itself is working properly and to satisfy our client's expectations [ 14]. 6. maintainance after deployment, we continued to provide ongoing support and maintenance to address any issues and incorporate necessary updates based on the client’s evolving needs [15]. iii. results and discussions our study shows that the system that we developed shows an excellent performance on the testing that we conducted by giving questionnaires and surveys. achieving 100% pass rate in the majority of its functional areas, such as managing the patients record, appointment scheduling and inventory tracking. these results highlight the system's effectiveness in meeting the clinic's core objectives, ensuring that essential tasks are streamlined and performed accurately. in terms of non-functional aspects, the system achieved impressive scores in key areas: reliability (97.14%), security (95.71%), and user experience (98.57%). these high scores emphasize the system’s robustness and ability to provide a secure, stable, and intuitive environment for both healthcare providers and patients. however, some minor areas for improvement were identified. specific functionalities scored 92.90%, and accessibility was rated at 90%. as shown in the figure 4 and figure 5 below is the total result summary of the functional and non-functional test result: vol.6, no.2, july 2025 | 105 figure 4. functional test summary figure 5. non-functional test summary on the other hand, the test result of the iso standard 25010 has also shown a great positive outcome, with improvements in functional suitability (96.14% to 96.47%), reliability (93.84% to 95.71%), and security (93.14% to 98.94%). attributes such as compatibility and interaction capability remained consistent, while slight declines were observed in performance efficiency (92.29% to 91.76%) and safety (91.58% to 91.11%). overall, the results highlight the system’s high compliance with iso standards, showcasing its strong functionality, security, and reliability. the figure 6 below is the iso standard 25010 result summary: vol.6, no.2, july 2025 | 106 figure 6. iso 25010 summary result iv. conclusions in conclusion, our study and the results of our testing and research shows that the system that we developed has greatly impacted the school clinic workflow, in which it allows the staff’s ability to efficiently manage the records of the patients, inventory, the health of the school premises. by this we can also conclude that the system not only helped the school clinic by managing the records, it also showed that by adding a feature such as integrating google meet in the system, leveraging their work progress in much efficient and precise manner. based on the feedback of our client, panels, and chairperson the following enhancements and recommendations are to be followed: 1. automated medical certificates: enable the system to generate medical certificates based on patient consultations 2. appointment reminders: implement automated reminders via email to ensure patients are notified of upcoming consultations. 3. improve security measures and system efficiency. 4. possibility of mobile application. 5. seamless integration of google meet 6. archiving or deletion of past medical record. acknowledgment we would like to give a heartfelt thanks to our very own advisers sir philipcris c. encarnacion, dit, sir syril glein t. flores, and to sir jonel e. ebol for giving us the opportunity to conduct this study, and by giving us the advice and the right path of having this research. to our family and friends thank you for the support. vol.6, no.2, july 2025 | 107 references [1] t. matingwina, "health, academic achievement and school-based," 2018. [2] mason stoltzfus , arshdeep kaur, avantika chawla, vasu gupta, f. n. u. anamika, and rohit jain, "the role of telemedicine in healthcare: an overview and update," the egyptian journal of, 2023. [3] pierre l. yong, robert s. saunders, and leighanne olsen (eds), the health care imperative: lowering cost and improving outcomes: workshop series summary, 2010. [4] m. e. holly m. satterfield, "technology use in health education: a review and future," the online journal of distance education and e-learning,, 2015. [5] n. al-shorbaji, improving healthcare access through digital health: the use of information and communication technologies, 2021. [6] f. c. dane, evaluating research: methodology for people who need to read research, sage publications, inc, 2011. [7] w. m. t. richard c. lee, uml and c++: a practical guide to object-oriented development second edition, united states: prentice-hall, inc., 2005. [8] manjunatha.n, "descriptive research," jetir (journal of emerging technologies and innovative research), vol. 6, no. 6, p. 865, 2019. [9] iso25000, "iso/iec 25010," 2020. [online]. available: https://iso25000.com/index.php/en/iso25000standards/iso-25010. [10] priyatna, b., rahman, t. k. a., hananto, a. l., hananto, a., & rahman, a. y. (2024). mobilenet backbone based approach for quality classification of straw mushrooms (volvariella volvacea) using convolutional neural networks (cnn). joiv: international journal on informatics visualization, 8(3-2), 1749-1754. [11] hananto, a. l., priyatna, b., & haris, a. (2020). application of prototype method on student monitoring system based on web. buana information technology and computer sciences (bit and cs), 1(1), 1-4. [12] ismail, d. a., huda, b., hilabi, s. s., & priyatna, b. (2024). penerapan desain ui/ux pada sistem penjualan berbasis web dengan metode desain thingking. innovative: journal of social science research, 4(2), 5737-5748. [13] susanto, s., priyatna, b., & permana, f. a. (2020). teacher monitoring application in teaching based on codeigniter framework in high schools. buana information technology and computer science, 1(1), 12-15. [14] hananto, a., pramono, e., & huda, b. (2022). application of recapitulation and staff performance assessment using standard working method. buana information technology and computer sciences (bit and cs), 3(1), 5-10. [15] novalia, e., na'am, j., nurcahyo, g. w., & voutama, a. (2020). website implementation with the monte carlo method as a media for predicting sales of cashier applications. systematics, 2(3), 118-131. vol. 6, no.1, january 2025 | 26 p-issn: 2715-2448 | e-issn: 2715-7199 buana information technology and computer sciences (bit and cs) design of interactive media for japanese m-learning based on android using uml and waterfall model apriade voutama1*, elfina novalia2 1*information systems, singaperbangsa university, karawang 2information systems, buana perjuangan university, karawang e-mail: apriade.voutama@staff.unsika.ac.id1*, elfinanovalia@ubpkarawang.ac.id2 received: 2024-12-18 | revised: 2024-12-29 | accepted: 2025-01-11 abstract the development of open-source technology has created many innovations created by developers in creating android-based applications. one of them is creating interactive media m-learning japanese language learning based on android. japanese m-learning is made to facilitate students in the learning process and provide interactive convenience in improving japanese language skills. m-learning is made using the android programming language with uml (unified modeling language) design analysis tools and the waterfall model. the initial analysis uses the usecase model to determine the actors involved in the system, and the flow of actor activities using the activity and sequence models and several other supporting models. the user interface is made to describe the system being built and the flowchart is designed to see the scheme running on the japanese m-learning interactive media system. the results of the analysis were carried out by taking satisfaction data from 30 student respondents, 81% stated that this interactive media was good. keywords: m-learning, japanese, uml, waterfall, android i. introduction the increasingly developing technology and information currently has a great influence on daily life for all groups [1]. especially in the development of smartphone technology. the use of various smartphones makes smartphones a primary need in communicating and obtaining information [2]. the development of smartphone devices that are increasingly high and relatively cheaper is the main factor in the use of smartphones in society. the digital marketing research institute emarketer in 2018 estimated that active smartphone users in indonesia reached more than 100 million people. with that number, indonesia will become the country with the fourth largest active smartphone users in the world after china, india, and america [3]. in the field of education, utilizing android smartphones has become a major part in supporting the needs and effectiveness of the teaching and learning process. the many innovations made in developing various educational applications make the learning process easy [4]. the many learning models that are built to answer the shortcomings of previous learning [5]. one of them is the interest of students who want to learn japanese, the expensive cost factor and educational places that are not all available make this a challenge in itself in the learning process of this interest. efforts need to be made, namely creating an interactive learning media in basic japanese language learning based on use via android smartphones [6]. so that with the existence of this interactive japanese language learning media, it can facilitate and help students interested in learning and increasing their knowledge in japanese. with the development that is already based on android, this learning media will greatly facilitate students interested in learning japanese without being hindered by space and time. the development of this interactive media cannot be separated from the planning tools and programming languages used. in the analysis and design process, the uml (unified modeling language) diagram model will be used, and the android programming language uses eclipse. vol.6 no.1, january 2025 vol. 6, no.1, january 2025 | 27 mobile learning is a new phenomenon with the effectiveness of countless tools in terms of education and high-quality learning process. smartphone is not just a phone but a device that can educate you at your will and at a time that is conveniently available to the user. mobile technology or smartphones in the digital era used in educational purposes are the core function of the spirit and development of mobile learning and referred to as ubiquitous learning [7]. m-learning in educational environment has brought several opportunities for learners and educators as indicated by the scope of m-learning that can be accessed anywhere only through android smartphones. m-learning helps learners in the learning process, join social media, find answers to their questions, enable team collaboration, facilitate knowledge sharing, and thus, improve their learning outcomes. several research studies are interested in developing m-learning applications in various contexts for different purposes with this interest driven by the availability of open-source development, accessibility of mobile infrastructure, mobile devices, low cost, learner motivation [8]. uml (unified modeling language) is one of the object-oriented modeling for designing and developing software or applications. uml modeling provides a standard for writing a blueprint system, which includes business concepts, writing classes in specific programming languages, components needed in a software system and database schema [9],[10]. uml is a very reliable model and is widely used by software/system developers because it is object-oriented [11]. this is because uml modeling provides a visual modeling language that allows system developers to create visuals in a standard form, easy to understand and equipped with an effective mechanism for sharing and communicating designs with others [12]. the basic concept of uml abstraction consists of three parts, namely structural classification, dynamic behavior, and model management and can be the main concepts as a reference when creating diagrams and views as categories of the diagram. uml consists of many models such as usecase diagrams, activity diagrams, sequence diagrams, class diagrams, statechart diagrams, collaboration diagrams, component diagrams, and deployment diagrams [13]. android is an operating system software used on mobile smartphone and tablet devices. the operating system is illustrated as a bridge between the device and its user, so that users can interact with their devices and run applications available on the device [14]. the use of android applications in the teaching and learning process is very important for both students and teachers. one of the functions of android applications for learning is the creation of applications related to student lessons where there are learning materials and a collection of practice questions that can be worked on by students. one of the android applications that can be developed for learning is the emodule application or learning module. android-based multimedia used in learning media makes learning more interesting than conventional information in the form of text by displaying animations, both audio and visual so that knowledge of historical information becomes more interesting [15]. the number of software/application developers in previous related studies such as research conducted by chandra et al. 2016, entitled "designing a mobile learning test of english for international communication (toeic) simulator application on android-based smartphones" where this study aims to create an android-based learning with a focus on learning and toeic simulation tests [16]. furthermore, research conducted by aisyiyah et al. 2019, entitled "mobile learning-based learning media for momentum and impulse material to improve students' critical thinking skills", this research was to create an android learning application related to momentum and impulse physics [17]. furthermore, research conducted by ibnu and wasis 2020, "development of mobile learningbased learning media for physical fitness activities of class x vocational high school students" this research produces android-based learning media for physical fitness learning containers for class x vocational high school students [18]. then previous research that utilized uml tools in the information system design process conducted by fifin and vina 2019, "utilization of uml (unified modeling language) in designing customer-to-customer type e-commerce information systems" research conducted utilizing uml tools as a system design model. vol. 6, no.1, january 2025 | 28 then research by suendri 2018, "implementation of uml (unified modeling language) diagrams in designing lecturer remuneration information systems with oracle database (case study: uin north sumatra medan)" the design carried out used uml diagrams in the design model. based on the many previous related studies, the design of this japanese language m-learning interactive media needs to be created. by utilizing the design tools, namely uml and the research model using the waterfall discipline to focus on working on each stage systematically. the following is uml consisting of 13 types of diagrams grouped into 3 categories [19]. the division of categories and types of diagrams can be seen in the following figure. figure 1. uml diagram figure 1 shows various diagrams that can be used to help design the system. in this study, several diagrams were used, such as usecase diagrams used to determine the actors and the overall system description, activity diagrams used to show the activities that actors can do in the system, sequence diagrams show detailed and more detailed descriptions of each actor's movement, class diagrams are used to determine interrelated classes and also as a concept of the database. then several other designs such as the design of the system interface that will be created and several flowchart designs to show the sequence from beginning to end. each stage in this study uses the waterfall discipline so that it can be more focused on each part of the design. so that with the innovation and development of the interactive media application of japanese m-learning, it can provide convenience for enthusiasts to increase their basic skills easily. ii. methods the research method used is using the computer science discipline sdlc (system development life cycle) or can be interpreted as a system development life cycle. in general, sldc has four stages, namely: planning, analysis, implementation, and maintenance [20]. these stages can be more detailed by utilizing a model from sdlc. this study uses an implementation model, namely waterfall, which is often referred to as the linear sequential model [21]. the waterfall model will start from the planning stage, then collect data from research case studies, and carry out the process of modeling the analysis results using uml, and carry out the design using uml, then enter development or coding, and the last stage is testing and implementation [22]. vol. 6, no.1, january 2025 | 29 figure 2. research framework iii. results and discussions a. analysis analysis is done by utilizing several uml models by producing an analysis design that will be made. the following is the definition of actors in the mobile learning application. table 1. definition of actor no actor description 1 user the user is the main actor in this mobile learning application. because this mobile learning application is offline, the user is the only actor involved in this system. 2 admin admin is an actor who acts as an operator in the mobile learning application. it is called an operator because the admin is an actor who inputs data into this application so that it can become an application that is ready to be used by the user. the following is the usecase plan that was formed. figure 3. usecase diagram activity diagrams describe how activities occur in the system to be designed. activity diagrams are the same as flowcharts that describe the processes that occur between actors and systems. admin and user activity diagrams are the steps of activities carried out by each actor in the application. the following are the admin and user activity diagrams. vol. 6, no.1, january 2025 | 30 figure 4. admin activity diagram figure 5. activity diagram user sequence diagrams are commonly used to describe scenarios or series of steps taken in response to an event to produce a specific output. data management sequence diagrams describe the activities carried out by the admin when inputting application data. figure 6. data management sequence diagram the indonesian tab menu sequence diagram illustrates the activities performed by the user after entering the main menu by selecting the indonesian tab. in this menu, the user can search for japanese words that are not understood. figure 7. sequence diagram of the indonesian menu tab the japanese tab menu sequence diagram illustrates the activities performed by the user after entering the main menu by selecting the japanese tab. in this menu, the user can search for the meaning of japanese words that are not understood. vol. 6, no.1, january 2025 | 31 figure 8. japanese tab menu sequence diagram the quiz menu sequence diagram illustrates the activities carried out by the user after entering the main menu by selecting the quiz menu. in this menu, the user can work on quizzes from previously available materials. figure 9. sequence diagram tab menu quiz the audio menu sequence diagram illustrates the activities performed by the user after entering the main menu by selecting the audio menu. in this menu, the user can listen to japanese language learning recordings from japanese nhk radio. figure 10. audio menu tab sequence diagram b. design at this design stage is a continuation of the analysis design results. some are used to describe the results of the analysis made previously. class diagram describes a collection of object classes. in this class diagram will be explained about the classes contained in the m-learning application. through the class diagram, the author will design the m-learning application by describing several classes that will be used in this m-learning application. vol. 6, no.1, january 2025 | 32 figure 11. class diagram interface design is done to describe in general how the application output will look. therefore, interface design can also be said to be a sketch of the application to be designed. interface design helps in creating application designs; therefore, it is very necessary to create an interface design first before designing the application design. figure 12. main menu & tab interface indonesia/japan c. system structure and flowchart at this stage, it shows part of the system structure and a flowchart of the process flow of each program activity. the menu structure is a general form of an application design to make it easier for users to run the application so that when running the application, users do not have difficulty in selecting the desired menus. the following is the form of the m-learning application interface menu structure. figure 13. m-learning application structure vol. 6, no.1, january 2025 | 33 the indonesian tab procedure is a procedure that occurs when the user selects the indonesian tab in the application, while the japanese tab procedure is a procedure that occurs when the user selects the japanese tab in the application. the flowchart of the indonesian and japanese tab procedural process flow can be seen in the following image. figure 14. flowchart of indonesian tab and japanese tab the grammar menu procedure is a procedure that occurs when the user selects the grammar menu in the application and the quiz menu procedure is a procedure that occurs when the user selects the quiz menu in the application. the grammar menu and quiz menu procedures can be seen in the following image. figure 15. flowchart of grammar tab and quiz tab d. implementation system implementation is a stage of system development. in this stage, several activities take place sequentially, namely starting from implementing the implementation plan, carrying out implementation activities, and implementation follow-up. mobile learning application testing can be done on all smartphones that have android os. in this menu, users can select japanese grammar materials that are available in the form of a list. this menu contains japanese grammar materials that can be used to learn the order or structure of sentences in japanese. vol. 6, no.1, january 2025 | 34 figure 16. gammar menu display the indonesian tab functions as a dictionary in searching for japanese words translated into indonesian. the indonesian tab display is in the form of a searching column, so users must input the first letter of the word they want to search for. after the loading process is complete, a list of words from the letters previously input will appear. figure 17. indonesian tab view the japanese tab functions as a dictionary in searching for indonesian words translated into japanese. the japanese tab display is in the form of a searching column, so users must input the first letter of the word they want to search for. after the loading process is complete, a list of words from the letters previously input will appear. vol. 6, no.1, january 2025 | 35 figure 18. japanese tab view in this menu, there are several nhk radio recordings that can be used as learning media in the form of sound. the display of this menu is in the form of a list of recordings, so that users can directly select which recording they want to listen to. in the detailed list display, there are three buttons that function to play audio, namely the play button (used to start audio), the pause button (used to stop audio temporarily) and the stop button (used if you want to end audio). figure 19. audio list view e. test analysis results the test results were conducted using a questionnaire distributed to 30 respondents after trying to use this interactive japanese m-learning application. where the contents of the questionnaire contain questions with three levels of answer choices, namely bad, enough, good. table 2. assessment questionnaire no questions bad enough good 1 how do you rate this android-based japanese language mlearning interactive media? 0 3 27 2 how do you rate the ease of using this system? 0 2 28 3 how do you rate the visual appearance of this japanese mlearning? 1 9 20 4 how do you rate the features in this japanese m-learning? 0 8 22 vol. 6, no.1, january 2025 | 36 5 how do you rate the speed or accuracy when this application runs? 0 5 25 based on table 2 above, the results of the rating levels given by respondents are obtained. the results are calculated and presented to obtain the following results. figure 20. assessment results figure 21. user feedback presentation from figure 20, it can be seen that the level of satisfaction obtained from this interactive japanese m-learning learning media. at the good assessment level = 81%, the enough level = 18%, while at the bad level = 1%. based on the assessment responses given by respondents, this application is worthy to be used and implemented for users who want to learn basic japanese. iv. conclusions the development of this japanese language m-learning interactive media learning has a very positive impact on supporting and providing convenience for students or enthusiasts who want to learn japanese. difficult places to study and human resources, then with this created application can help and the process of developing knowledge in learning japanese. based on the results of the analysis, 81% were obtained with a good level of satisfaction so that the m-learning application is useful and provides convenience. so that with the existence of the japanese language mobile learning application, it can be more efficient in time in searching and purchasing japanese language books and can also save expenses because it maximizes the use of smartphones as one of the technological developments. in further developments, this m-learning application is still worth developing by adding new features so that this learning media is more complex and by developing more attractive and sweet interface displays so that this learning media can be implemented sustainably. vol. 6, no.1, january 2025 | 37 references [1] s. adhi, k. marhadini, i. akhlis, and i. sumpono, “pengembangan media pembelajaran berbasis android pada materi gerak parabola untuk siswa sma,” upej unnes phys. educ. j., vol. 6, no. 3, pp. 38–43, 2017, doi: 10.15294/upej.v6i3.19315. [2] a. voutama, i. maulana, and n. ade, “interactive m-learning design innovation using android-based adobe flash at wfh (work from home),” sci. j. informatics, vol. 8, no. 1, pp. 127–136, 2021, doi: 10.15294/sji.v8i1.27880. [3] anita adesti and siti nurkholimah, “pengembangan media pembelajaran berbasis android menggunakan aplikasi adobe flash cs 6 pada mata pelajaran sosiologi,” edutainment, vol. 8, no. 1, pp. 27–38, 2020, doi: 10.35438/e.v8i1.221. [4] a. voutama, “perancangan aplikasi m-discussion berbasis android sebagai wadah diskusi sekolah,” syntax j. inform., vol. 7, no. 2, pp. 116–124, 2018. [5] e. novalia et al., “website implementation with the monte carlo method as a media for predicting sales of cashier applications,” vol. 2, no. 3, pp. 118–131, 2020. [6] a. aditya and d. w. s. susanto, “rancang bangun aplikasi media pembelajaran bagi siswa penyandang tuna rungu berbasis android,” techno.com, vol. 20, no. 4, pp. 540–551, 2021, doi: 10.33633/tc.v20i4.5216. [7] m. i. qureshi, n. khan, s. m. ahmad hassan gillani, and h. raza, “a systematic review of past decade of mobile learning: what we learned and where to go,” int. j. interact. mob. technol., vol. 14, no. 6, pp. 67–81, 2020, doi: 10.3991/ijim.v14i06.13479. [8] m. al-emran, v. mezhuyev, a. kamaludin, and m. alsinani, “development of m-learning application based on knowledge management processes,” acm int. conf. proceeding ser., pp. 248–253, 2018, doi: 10.1145/3185089.3185120. [9] f.sonata, “pemanfaatan uml (unified modeling language) dalam perancangan sistem informasi e-commerce jenis customer-to-customer,” j. komunika j. komunikasi, media dan inform., vol. 8, no. 1, p. 22, 2019, doi: 10.31504/komunika.v8i1.1832. [10] a. rohmat, b. a. dermawan, a. voutama, and b. gunadi, “sistem pakar penentuan jenis budidaya ikan air tawar berdasarkan lokasi dan kualitas air,” j. teknol. dan inf., vol. 11, no. 2, pp. 96–110, 2021, doi: 10.34010/jati.v11i2.3490. [11] b. adhi pamungkas, a. voutama, and b. nurina sari, “sistem pakar deteksi dini hiv/aids dengan metode forward chaining dan certainty factor expert system of hiv/aids early detection with forward chaining and certainty factor method,” j. inf. technol. comput. sci., vol. 4, no. 1, 2021. [12] a. voutama and e. novalia, “perancangan aplikasi m-magazine berbasis android sebagai sarana mading sekolah menengah atas,” j. tekno kompak, vol. 15, no. 1, p. 104, 2021, doi: 10.33365/jtk.v15i1.920. [13] suendri, “implementasi diagram uml (unified modelling language) pada perancangan sistem informasi remunerasi dosen dengan database oracle (studi kasus: uin sumatera utara medan),” j. ilmu komput. dan inform., vol. 3, no. 1, pp. 1–9, 2018, [online]. available: http://jurnal.uinsu.ac.id/index.php/algoritma/article/download/3148/1871. [14] t. wiranda and m. adri, “rancang bangun aplikasi modul pembelajaran teknologi wan berbasis android,” voteteknika (vocational tek. elektron. dan inform., vol. 7, no. 4, p. 85, 2020, doi: 10.24036/voteteknika.v7i4.106472. [15] f. tahel and e. ginting, “perancangan aplikasi media pembelajaran pengenalan pahlawan nasional untuk meningkatkan rasa nasionalis berbasis android,” teknomatika, vol. 09, no. 02, pp. 113–120, 2019. [16] y. f. chandra, n. dwiyani, and y. huda, “perancangan aplikasi mobile learning test of english for international communication (toeic) simulation pada smartphone berbasis android,” voteteknika (vocational tek. elektron. dan inform., vol. 5, no. 1, 2017, doi: 10.24036/voteteknika.v5i1.6167. [17] a. h. ngurahrai, s. d. farmaryanti, and n. nurhidayati, “media pembelajaran materi momentum dan impuls berbasis mobile learning untuk meningkatkan kemampuan berpikir kritis siswa,” berk. ilm. pendidik. fis., vol. 7, no. 1, p. 62, 2019, doi: 10.20527/bipf.v7i1.5440. [18] i. a. pamungkas and w. d. dwiyogo, “pengembangan media pembelajaran berbasis mobile learning untuk aktifitas kesegaran jasmani siswa kelas x sekolah menengah kejuruhan,” vol. 6, no.1, january 2025 | 38 sport sci. heal., vol. 2, no. 5, pp. 272–278, 2022, doi: 10.17977/um062v2i52020p272-278. [19] a. voutama, “sistem antrian cucian mobil berbasis website menggunakan konsep crm dan penerapan uml,” komputika j. sist. komput., vol. 11, no. 1, pp. 102–111, 2022, doi: 10.34010/komputika.v11i1.4677. [20] a. t. j. harjanta and b. a. herlambang, “rancang bangun game edukasi pemilihan gubernur jateng berbasis android dengan model addie,” j. transform., vol. 16, no. 1, p. 91, 2018, doi: 10.26623/transformatika.v16i1.894. [21] d. zaliluddin and r. rohmat, “perancangan sistem informasi penjualan berbasis web (studi kasus pada newbiestore),” infotech j., vol. 4, no. 1, p. 236615, 2018. [22] agariadne dwinggo samala, bayu ramadhani fajri, and fadhli ranuharja, “desain dan implementasi media pembelajaran berbasis mobile learning menggunakan moodle mobile app,” j. teknol. inf. dan pendidik., vol. 12, no. 2, pp. 13–20, 2019, [online]. available: http://tip.ppj.unp.ac.id. vol. 5, no.1 januari 2024 | 1 j.valarmathi1, v.t.kruthika2 implementation of marker-based tracking method on augmented reality in multimedia learning (case study of stmik tegal) gunawan gunawan1*, wresti andriani2, sawaviyya anandianskha3, muhammad indratama4 1,2,3,4 stmik ymi tegal, jl. pendidikan no. 1 kota tegal 1 gunawan.gayo@gmail.com, 2 wresty.adriani@gmail.com, 3 sawaviyyaa@gmail.com, 4 mohammadindratama@gmail.com abstract introducing campus locations for new students or address seekers is an important activity. multimedia learning is not only a tool for creating harmonious presentations and alternatives that combine visual and audio media; technology can be used for its tools. augmented reality (ar) is one of them. augmented reality is helpful as a combination of virtual and reality devices that operate interactively in a realtime natural environment. based marker tracking is a method used to make objects into two dimensions and three dimensions whose process begins with directing the marking object by the user using the camera on the mobile device until the camera reads the object. light intensity affects detection success, and distance calculation also becomes essential. if the marker is successfully detected, the application will convert it into a 3-dimensional object as the final result. in this study, a location search will be carried out for the stmik tegal campus building using augmented reality based on the based marker tracking method to produce the most ideal conditions to be able to display 3d objects from the stmik tegal building, which is a distance of 15 to 25 cm with bright light using android, so that this application can be used to find the location of the stmik tegal building. keywords: augmented reality, based marker tracking, multimedia, stmik tegal i. introduction multimedia learning at this time is beneficial for both business and educational fields [1]. in the business field, multimedia can be used in various media profiles, such as product and company profiles. some even make it as an information stall (kiosk) and learn in an online learning system [2]. while in the field of multimedia education, this can be used as a teaching medium either in the classroom or independently or self-taught[3]. in everyday life, it can be used as a tool to find the location of a place. multimedia [4] is an application from a computer whose presentation combines text, sound, images, animation, audio, and video with tools and links from the computer so that users can interact, navigate, create, and communicate [5]. so, multimedia can also be a presentation tool [6]. using digital media intermediaries such as mobile applications, multimedia learning offers users various options, enabling technology to make human life easier [7]. augmented reality (ar) is a technology that combines two-dimensional (2d) and threedimensional (3) virtual objects in a three-dimensional (3d) natural environment and then projects these virtual objects in real time [[8]. until now, many applications have adopted this ar technology as a game, business and educational media. as on smartphones that are capable of displaying 3d objects that are informed, so they do not get bored when used [9], [10]. marker based tracking is augmented reality (ar) [11], [12]. p-issn : 2715-2448 | e-issn : 2715-7199 vol.5 no.1 januari 2024 buana information technology and computer sciences (bit and cs) mailto:gunawan.gayo@gmail.com mailto:wresty.adriani@gmail.com mailto:sawaviyyaa@gmail.com mailto:mohammadindratama@gmail.com vol. 5, no.1 januari 2024 | 2 in a previous study, [13] which conducted research on augmented reality-based distance learning multimedia in the midst of a pandemic, the results of which stated that using augmented reality-based multimedia can increase students' science literacy. likewise [14], [15]. in another study that has been conducted [16] using the interactive system multimedia design and development (ismdd) method on augmented reality, the introduction of the uki maluku building, the results of which are based on tests that have been carried out, have met the standards with a good percentage of success. based on the background above, this research will use augmented reality technology with the marked based method to find the location of the tegal stmik building the selection of this method is expected to facilitate users in finding locations in the context of introducing the tegal stmik building. in its application, the operating system on android will be used so that it is more flexible and can be used anytime and anywhere. markers will be in the form of posters, to facilitate the process of recognizing and distributing this technology to users. this poster will be detected through an android device to display a three-dimensional animated object. ii. methods the method to be used in this research is the multimedia development life cycle (mdll) which is suitable for the development of multimedia-based systems which have six stages, namely concept, design, material collecting, assembly, testing and distribution [17]. these stages are shown in figure 1. figure 1. diagram mdll based on figure 1, the six stages include: 1. concept this stage begins with determining the purpose of making the application, the target application users and what materials will be used or displayed. 2. design at this stage it aims to make detailed specifications about the project architecture, appearance and material needs. 3. material collecting then at this stage is to collect material according to what has been determined or set at the design stage above. 4. assembly then at this assembly stage will then be made applications based on the design stage, on the results of information obtained at the material collection stage, using programming software such as unity 3d. vol. 5, no.1 januari 2024 | 3 5. testing at this stage, it aims to ensure that the results of making the application are in accordance with what has been designed. at this stage, the blackbox method is also carried out on the user interface, by ensuring the accuracy of the model on the marker, button functions and animations obtained. if there is a failure or bug, improvements or revisions must be made. 6. distribution this stage is carried out when testing has been completed and declared suitable for use, then this stage will be disseminated so that this application can be used by users. flow diagram in this study as a material used to describe the sequence of processes in detail with relationships in other processes in a program [18]. the flowchart of augmented reality is shown in figure 2. figure 2. diagram alur augmented reality 7. android is the operating system used is the development of the linux system. android itself was developed by the star up with the same name android inc, in 2005. google bought android and took over the developer's job [19]. 8. unity is a cross-platform game development application, unity 3d is a tool for creating 3d [20]. unity 3d can generate terrain, import a model of the main building that has been built into the terrain and placed according to the actual angle. unity has an sdk provided by quantum to help developers create ar applications [21]. computer vision to recognize and track planar images (image targets and simple 3d objects such as squares in realtime [22]. iii. results and discussions in this study, which uses the augmented reality method, it is carried out as a search and location recognition tool for the stmik tegal building, the tool that will be used in this study is as in table 1. vol. 5, no.1 januari 2024 | 4 table 1. software and hardware used kebutuhan hardware kebutuhan software 1. laptop (asus x555ba), processor (intel core i7 gen 10th), memory (16gb ddr4) dan hard disk (1000 gb) 2. smartphone (oppo a31) 3. printer (epson l1110) 1. sistem operasi (windows 10 pro 64 bit) 2. blender 3d 3. unity 3d 4. adobe photoshop 5. sdk (vuforia dan android) and airdroid the results of the six steps in the mdlc consisting of 6 stages have been carried out: 1. concept the concept produced at this stage is the purpose of the application, namely the existence of media that can display information about the location of the stmik tegal building building. applications used by the android operating system developed with programming languages on the unity engine, namely c + programming language. 2. researchers at this stage make designs consisting of system architecture, flowcharts, storyboards and interface designs. the architecture of the system is shown in figure 2 figure 3. system architecture in figure 3 it appears that the application will run on devices with the android operating system, then this application will use the camera as a marker scanner media, if this marker can be read on the camera to be displayed, namely 3d objects on the layer of android. while the skyboard design that will be made as in figure 4. vol. 5, no.1 januari 2024 | 5 figure 4. storyboard design the interface design that will be used in this study is designed as in figure 5. next, here are the components of the functional picture in a system. figure 5. use case diagram of augmented reality with unity app figure 5 shows the design of a use case diagram for augmented reality, the results of the user interface design are shown in figure 6. figure 6. user interface at the material collecting stage, a needs analysis and collection of materials will be used in application development. material in the form of information around the address and location and vol. 5, no.1 januari 2024 | 6 appearance of the stmik tegal building which will be displayed in the augmented reality application the application development process will use several supporting applications such as designing photoshop application assets because it is easier to use while for the formation of 3d objects using archicad 23 because it can display the design and appearance of buildings and blender applications for texture and effects and 3d building object light. the unity 3d application is also used to perform camera settings, database connections and create interface displays. on the main menu unterface there is a button for the stmik tegal building, the building button or location before the stmik building and the building after the location of the stmik tegal gedaung then the exit button to close the application. in figure 6, you can see the location of the tegal stmik building seen from the air which illustrates the entire location around the tegal stmik building. the tester stage is carried out by adjusting the scan distance, light testing distance and testing according to the version of android owned by use, along with the test results in the table. figure 7. environmental location of stmik tegal building in picture 6, you can see that stmik tegal building is located next to saphire housing (left) and next to man building (green). the tester stage is carried out by adjusting the scan distance, light testing distance and testing according to the version of android owned by use, following the test results in table 2. table 2. distance testing no distance (cm) status 1 5 the building is not visible 2 15 building view 3 25 building view 4 50 the building began to become invisible 5 100 the building is not visible based on application testing based on light as table 3. table 3. light testing no light status 1 bright building view 2 dark the building is not visible testing on android performed on this version is shown in table 4. vol. 5, no.1 januari 2024 | 7 table 4. android versions for testing no versi android status 1 xiaomi redmi plus the application runs 2 samsung galaksi a71 the application runs 3 oppo a31 the application runs iv. conclusions this researched augmented reality technology can be implemented as a tool or medium for recognition or location search of gedang stmik tegal by displaying 3d objects from the stmik tegal building. the method used in this application is marked based tracker used by printing markers and inserting them in the prosur can be the easiest way for prospective new students to find out about information through brochures and can see the visualization from the stmik tegal building, from the test results carried out the most ideal conditions to be able to display 3d objects from the tegal stmik building which is a distance of 15 to 25 cm with bright light using android. references [1] i. listiani, “analisis pentingnya sistem informasi manajemen dalam teknologi informasi dan komunikasi saat ini,” informasi, teknol. dan komun., vol. 1, pp. 1–15, 2021. [2] c. bino evin and a. didimus rumpak, “analisis pengembangan desain grafis dalam aplikasi photoshop sebagai peluang bisnis mahasiswa institut bisnis dan multimedia asmi,” j. sist. inf., vol. 1, no. 2, pp. 33–40, 2019. [3] p. w. aditama, i. nyoman widhi adnyana, and k. ayu ariningsih, “augmented reality dalam multimedia pembelajaran,” pros. semin. nas. desain dan arsit., vol. 2, pp. 176–182, 2021. [4] p. manurung, “multimedia interaktif sebagai media pembelajaran pada masa pandemi covid 19,” al-fikru j. ilm., vol. 14, no. 1, pp. 1–12, 2021, doi: 10.51672/alfikru.v14i1.33. [5] n. amalia, “aplikasi flash player berbasis multimedia interaktif menggunakan adobe reader,” edutech j. ilmu pendidik. dan ilmu sos., vol. 7, no. 2, 2021. [6] n. a. suryandaru, “penerapan multimedia dalam pembelajaran yang efektif,” j. pendidik. dan pengajaran guru sekol. dasar, vol. 3, pp. 88–91, 2020. [7] r. sumarlin, mario, d. n. anggraini, and d. hidayat, “review dan analisis multimedia learning berbasis cerita rakyat sunda melalui mobile apps,” j. demandia, vol. vol. 07 no, 2022, doi: https://doi.org/10.25124/demandia.v7i2.4404. [8] t. abdulghani and b. p. sati, “pengenalan rumah adat indonesia menggunakan teknologi augmented reality dengan metode marker based tracking sebagai media pembelajaran,” media j. inform., vol. 11, no.1, pp. 43–50, 2019, doi: https://doi.org/10.35194/mji.v11i1.770. [9] i. hastuti and a. n. azura, “dan pahlawan lokal banjarmasin berbasis mobile augmented reality pada museum wasaka banjarmasin,” vol. 14, no. 1, pp. 18–27, 2022. [10] nopriyanti and p. sudira, “pengembangan multimedia pembelajaran interaktif kompetensi dasar pemasangan sistem penerangan dan wiring kelistrikan di smk,” j. pendidik. vokasi, vol. 5, 2015. [11] m. santoso, c. r. sari, and s. jalal, “promosi kampus berbasis augmented reality,” j. edukasi elektro, vol. 5, no. 2, pp. 105–110, 2021, doi: 10.21831/jee.v5i2.43496. [12] a. a. aldriyan and s. amini, “penerapan metode marker based tracking untuk pembelajaran anak berkebutuhan khusus,” skanika, vol. 3, no. 4, pp. 1–6, 2020. [13] m. ahied, l. k. muharrami, a. fikriyah, and i. rosidi, “improving students’ scientific literacy through distance learning with augmented reality-based multimedia amid the covid-19 pandemic,” j. pendidik. ipa indones., vol. 9, no. 4, pp. 499–511, 2020, doi: 10.15294/jpii.v9i4.26123. [14] h. bursali and r. m. yilmaz, “effect of augmented reality applications on secondary school students’ reading comprehension and learning permanency,” comput. human behav., vol. 95, no. june 2018, pp. 126–135, 2019, doi: 10.1016/j.chb.2019.01.035. [15] j. garzón, kinshuk, s. baldiris, j. gutiérrez, and j. pavón, “how do pedagogical approaches vol. 5, no.1 januari 2024 | 8 affect the impact of augmented reality on education? a meta-analysis and research synthesis,” educ. res. rev., vol. 31, p. 100334, 2020, doi: 10.1016/j.edurev.2020.100334. [16] r. siwalette and h. tuhuteru, “augmented reality untuk pengenalan gedung universitas kristen indonesia maluku,” jipi (jurnal ilm. penelit. dan pembelajaran inform., vol. 7, no. 4, pp. 1411–1417, 2022, doi: 10.29100/jipi.v7i4.3927. [17] s. alisyafiq, b. hardiyana, and r. p. dhaniawaty, “implementasi multimedia development life cycle pada aplikasi pembelajaran multimedia interaktif algoritma dan pemrograman dasar untuk mahasiswa berkebutuhan khusus berbasis android,” j. pendidik. kebutuhan khusus, vol. 5, no. 2, pp. 135–143, 2021, doi: 10.24036/jpkk.v5i2.594. [18] s. a. olii, t. abdillah, and r. takdir, “augmented reality katalog penjualan perangkat keras menggunakan magic book berbasis android,” diffus. j. syst. …, vol. 1, no. 2, 2021. [19] a. belkhir, m. abdellatif, r. tighilt, n. moha, y. g. gueheneuc, and e. beaudry, “an observational study on the state of rest api uses in android mobile applications,” proc. 2019 ieee/acm 6th int. conf. mob. softw. eng. syst. mobilesoft 2019, pp. 66–75, 2019, doi: 10.1109/mobilesoft.2019.00020. [20] aditya fajar ramadhan, ade dwi putra, and ade surahman, “aplikasi pengenalan perangkat keras komputer berbasis android menggunakanaugmented reality (ar),” j. teknol. dan sist. inf., vol. 2, no. 2, pp. 24–31, 2021. [21] n. rianto, a. sucipto, and r. dedi gunawan, “pengenalan alat musik tradisional lampung menggunakan augmented reality berbasis android (studi kasus: sdn 1 rangai tri tunggal lampung selatan),” j. inform. dan rekayasa perangkat lunak, vol. 2, no. 1, pp. 64–72, 2021. [22] k. g. nalbant and ş. uyanik, “computer vision in the metaverse,” j. metaverse, vol. 1, no. 1, pp. 9–12, 2021. vol. 6, no.1, january 2025 | 1 multiagent systems as an approach to building fuzzy voter communities using fuzzy languages michel milambu belangany1*, pièrre kasengedia motumbe2, eugène mbuyi mukendi3 1,3 department of mathematics and computer science, university of kinshasa, 2department of engineering computing, higher institute of applied techniques email: michel.milambu@unikin.ac.cd, pierre.kasengedia@ista.ac.cd, eugenembuyi@gmail.com received: 2024-04-12 | revised: 2024-12-26 | accepted: 2025-01-10 abstract this paper explores the use of fuzzy set theory to model the behavior of voters in a multi-agent electoral environment. voters, represented as fuzzy agents, communicate using imprecise language to form communities based on shared linguistic terms. by leveraging graph theory, we construct a model of a fuzzy voting system where agents are linked based on the similarity of their fuzzy language. the proposed approach focuses on identifying, constructing, and extracting communities of fuzzy voters without delving into their relational dynamics. using fuzzy set membership functions, we define linguistic variables that reflect the imprecision in voter behavior. the study introduces an algorithm to detect communities by creating links between fuzzy voters, ultimately forming groups based on their linguistic similarities. results demonstrate that fuzzy communities can be successfully constructed, where the membership function quantifies the degree of belonging of voters to specific communities. this method contributes to a better understanding of voting behavior in complex, heterogeneous systems and offers a novel approach to community detection in multi-agent systems. keywords : fuzzy language, multiagent systems, community, construction and voter i. introduction in this article, a community is formed of voters or agents who speak the same language, in our case it is the fuzzy language that concerns us to create links and form community of fuzzy voters. referring to the article on taking into account imprecision in the modeling voting in a multi-agent environment [1], in which he defines an electoral system is a set of individuals considered as agents in a multi-agent system in which voters communicate with each other and with the environment. in such a system, it is often difficult to understand the behavior of an agent that we call a voter. this is why, in this paper, we use fuzzy set theory as an approach to model the behavior of an imprecise voter in an electoral environment. it will be just a question of presenting a model of a voter with fuzzy behavior using mathematical approaches in this environment considered as a multi-agent environment and to propose the algorithms as the tools of computer modeling [1], there is reason to identify or build community by creating links between vague voters based on their languages without worrying about their relationships. some authors have used the term community, multi-agent system in particular: complex networks have a large number of nodes and edges, which prevents the understanding of network structure and the discovery of valid information. this paper proposes a new community detection method for simplified networks. first, a similarity measure is defined, the path and attribute information can reflect the potential relationship between nodes that are not directly connected [2]. community detection aims to discover hidden community or groups in complex networks and is essentially unsupervised clustering behavior. however, most of the existing unsupervised methods are designed for homogeneous networks; therefore, they cannot effectively handle heterogeneous structures and rich semantic information. under such a situation, it is difficult to accurately detect community in heterogeneous networks that better reflect the real world [3]. a community is a set of nodes in a network where the density of connections is high [4]. p-issn : 2715-2448 | e-issn : 2715-7199 vol.6 no.1 january 2025 buana information technology and computer sciences (bit and cs) mailto:michel.milambu@unikin.ac.cd mailto:eugenembuyi@gmail.com vol. 6, no.1, january 2025 | 2 using this definition on the real world around us, we can confirm that an electoral system is a multiagent system. the voting system employed by several countries or states is considerably a typical example of a multi-agent system. indeed, we can specify the set that characterizes this system. in this paper, the consensus problem of heterogeneous multi-agent systems under directed topology is investigated [3],[4]. specifically, this system is composed of three classes of agents respectively described by first-order, second-order and third-order integrator dynamics [5],[6]. by the aid of linear filter, graph theory and matrix theory, the consensus problem is realized based on the two proposed consensus protocols [7],[8]. moreover, group consensus can also be solved by adjusting parameters. this theory applicated in simulating spatiotemporal dynamics of urban underground space development using multi-agent system: a case study in changzhou city, china [9], [10]. consensus-based distributed connectivity control in multi-agent systems, in this paper it present distributed connectivity control problem in networked multi-agent systems [11],[6]. the system communication topology is controlled through the algebraic connectivity measure, the second smallest eigenvalue of the communication graph laplacian [12],[13]. the algebraic connectivity is estimated locally in a decentralized manner through a trust based consensus algorithm, in which the agents communicate the perceived quality of the communication links in the system with their set of neighbors [14],[15]. ii. methods in our article, we will use this notion from graph theory to allow us to create links between imprecise voters which we otherwise call fuzzy voters. thus, a graph g is made up of two sets [2] : 1. a set x = (x1, x2, …, xn) of elements called vertices or nodes, materialized by points: 2. a set u = (u1, u2, ..., un) of ordered pairs (i,j) with i  x and j  x. the elements of this set are called “arcs” or “branches” a graph is therefore noted g = (x, u). but, in this article we will note c = (e, l): 1. a set e = (e1, e2, …, en) elements called fuzzy voters or agents, materialized by agents: 2. a set l = (l1, l2, ..., lm) of ordered pairs (i,j) with i  e and j  e. the elements of this set are called “links” or “relations”. we use the theory of fuzzy subsets which will allow us to present the imprecise behavior of a voter in an electoral system. let x be a reference set and let x be any element of x. a fuzzy set a of x is defined as the set of couples (milambu, kafunda & mbuyi, 2024) : 𝐴 = {(𝑥, 𝜇𝐴(𝑥)), 𝑥 ∈ x} (1) where : 𝜇𝐴: 𝑋 → [0, 1] (2) thus, a fuzzy set a of x is characterized by a membership function that associates, to each element x of x a real in the interval [0, 1]; 𝜇𝐴(𝑥) represents the degree of membership of x to a. thus, the closer the value of 𝜇𝐴(𝑥) is to unity, the higher the degree of membership of x to a [16]. if we have : 𝜇𝐴: 𝑋 → {0, 1 } we find the boolean case: either x belongs to 𝐴(𝜇𝐴 = 1) or it does not belong to 𝐴(𝜇𝐴 = 0). and the following case is very useful in the sense that an element belongs partially: let x belong partially to 𝐴(0 < 𝜇𝐴(𝑥) < 1) it is important to specify that the fuzzy set is considered as empty if the membership degrees of all the elements of the universe are all equal to zero. 𝐴 = ∅ ⇔ 𝜇𝐴 (𝑥) = 0, ∀𝑥 ∈ 𝑋 (3) vol. 6, no.1, january 2025 | 3 two fuzzy sets are equal if their membership degrees are equal for all elements of the reference set, i.e., if both fuzzy sets have the same membership function [17]. two fuzzy sets a and b, defined on the same reference set x are equal if: a = b ⇔ 𝜇𝐴(𝑥) = 𝜇𝐵(𝑥), ∀𝑥 ∈ 𝑋 (4) inclusion 𝐴 ⊆ 𝐵 ⟺ ∀𝑥 ∈ 𝐸, 𝜇𝐴(𝑥) ≤ 𝜇𝐵(𝑥) (5) union 𝐴 ∪ 𝐵 = max (𝜇𝐴(𝑥), 𝜇𝐵(𝑥))), ∀x ∈ e (6) intersection 𝐴 ∩ 𝐵 = min (𝜇𝐴(𝑥), 𝜇𝐵(𝑥))), ∀x ∈ e (7) �̅� ∶ ̅ is said to be complementary to a if its membership function satisfies: 𝜇�̅�(𝑥) = 1 − 𝜇𝐴(𝑥), ∀𝑥 ∈ 𝐸 (8) the reference set of a natural language word is called the discourse universe. the one-word discourse universe is a set of terms that evoke the same concept but to different degrees. it may or may not be finished [18]. a. linguistic variable a linguistic variable represents a state in the system to be adjusted. each linguistic variable is characterized by a set such that: {𝑣, 𝐸(𝑣), 𝑈, 𝑅, 𝑆} or : v:is the name of the variable e(v): is the set of linguistic values that v can take u:is the universe of discourse associated with the base value r: is the syntactic rule to generate the linguistic values of v s: is the semantic rule to associate a meaning with each linguistic value [19]. for the case of this thesis, the model is as follows: the linguistic variable v = candidate choice this variable can be defined with a set of terms 1) e(v) = {good, very good, extremely good, not very good, bad, very bad, extremely bad, little bad}: which form his universe of discourse 2) u = [0%, 100%] 3) the basic value is the choice of the candidate 4) the term “good” represents a linguistic value it can be interpreted as: “choices greater than 50%” “choices smaller than 50%” b. linguistics of a voter below we present some vague words from voters: 1) we will see; 2) i could vote good candidate; 3) i see x doing but y also sometimes z; 4) i can vote x good! we'll see because, yes too but z had done well in all ways i don't know yet who to vote for. 5) i'm not interested in this yet 6) i want to see first [20],[28]. 7) etc… vol. 6, no.1, january 2025 | 4 𝐿𝑎𝑛𝑔𝑒𝑖 = {𝑚1, 𝑚2, … , 𝑚𝑝} with : 𝐿𝑎𝑛𝑔𝑒𝑖 : 𝑣𝑜𝑡𝑒𝑟 𝑙𝑎𝑛𝑔𝑢𝑎𝑔𝑒 𝑚𝑖 : words or terms used by a voter c. algorithm proposal table 1. algorithm proposal algorithm proposal 1. parameter 𝐿𝑎𝑛𝑔𝑒𝑖 = {𝑚𝑖} and 𝐿𝑎𝑛𝑔𝑒𝑗 = {𝑚𝑗} / i=1,2,…, n and j=1,2,…,p 𝑉𝐿𝑖𝑛𝑔𝑢𝑖𝑠𝑡𝑖𝑐𝑠 = 𝑐ℎ𝑜𝑖𝑐𝑒 , 𝑇(𝑐ℎ𝑜𝑖𝑐𝑒) = {𝑡1, 𝑡2, … , 𝑡𝑛} 2. output 𝐶 = (𝐸, 𝐿) / c is community, e is voters set and l is link 3. repeate 4. for i = 1 to n do 5. for j = 1 to p do 6. 𝐸 = {𝑒𝑖, 𝑒𝑗 ∶ 𝑒𝑖 𝑎𝑛𝑑 𝑒𝑗 𝑖𝑠 𝑣𝑜𝑡𝑒𝑟𝑠} 7. if 𝜇𝐿(𝑙(𝑒𝑖)) = 𝜇𝐿 (𝑙(𝑒𝑗)) , ∀𝑙(𝑒𝑖), 𝑙(𝑒𝑗) ∈ u then 8. create the link between 𝑒𝑖 𝑎𝑛𝑑 𝑒𝑗 9. 𝐿 = {𝑒𝑖𝑒𝑗 ∈ 𝐸𝑥𝐸} 10. affect them in 𝐶𝑘 / 11. else no link 12. end if 13. end. table 2. the matrix of decision 𝑽𝑳𝒆𝒋 . 𝒕𝒋 𝑽𝑳𝒆𝒋+𝟏 𝒕𝒋+𝟏 𝑽𝑳𝒆𝒊 . 𝒕𝒊 vf vf 𝑽𝑳𝒆𝒊+𝟏 . 𝒕𝒊+𝟏 vf vf this table is a matrix representing the links between voters with fuzzy language. the link is only possible between voters if these voters use vague terms about their choices [21],[22],[23],[24]. iii. results and discussions we thus open the discussions by presenting the different results obtained on the construction of community based on fuzzy voters. a community built from fuzzy voters is also fuzzy. figure 1. identify of community vol. 6, no.1, january 2025 | 5 the identification of community of voters with the same imprecise language in the choice of candidates in a population of voters. figure 2. extract of voters groups the fuzzy voters are grouped according to whether they use the fuzzy terms in order to build community and the other ungrouped ones do not interest us because they have a precise choice [26],27]. figure 3. community trained as we can see in the figure above, the extraction of fuzzy community from fuzzy voters. the membership function gives a value of 0.3 for a community of fuzzy voters whose language revolves around percentage 30 to 40. vol. 6, no.1, january 2025 | 6 figure 5. degree of belonging to a vague community we present a figure or diagram of fuzzyfication and defuzzyfication from classical language to fuzzy language and from fuzzy to classical language below. in this diagram, we have as input the classical language which is fuzzyfied taking into account the linguistic variable and all the terms associated with the different fuzzy rules. figure 6. fuzzyfication and defuzzyfication scheme figure 7. degree of belonging of fuzzy term. figure 8. degree of belonging two fuzzy terms vol. 6, no.1, january 2025 | 7 table 2. fuzzy matrix cognitive choice bad good choice bad b g good g g iv. conclusion this article is a continuation of the publication on taking into account imprecision in a voter's behavior. it was therefore a question of this article proposing an approach for constructing community of fuzzy voters based on the terms used in the language of voters. throughout this article, we have used agents to represent the community of these so-called vague voters. we focused on the identify, construction and extraction of fuzzy community as presented in the different figures of this article. the use of this approach makes it possible to construct groups of voters without seeking to know the relationships between voters, but only exploit their languages in order to identify and construct communities based on the proposed algorithm. references [1] m. milambu, p. kafunda and e. mbuyi, taking into account imprecision in the modeling voter in a multi-agent environment, p-issn : 2715-2448 | e-issn : 2715-7199 vol.5 no.1 januari 2024, bit and cs. [2] h. zheng, h. zhao and g. ahmadi, towards improving community detection in complex networks using influential nodes, journal of complex networks, volume 12, issue 1, february 2024, cnae001, https://doi.org/10.1093/comnet/cnae001 [3] y. zheng, q. zhao, j. ma, and l. wang, “systems & control letters second-order consensus of hybrid multi-agent systems ✩,” syst. control lett., vol. 125, pp. 51–58, 2019, doi: 10.1016/j.sysconle.2019.01.009. [4] h. geng, h. wu, j. miao, s. hou, and z. chen, “consensus of heterogeneous multi-agent systems under directed topology,” ieee access, vol. 10, pp. 5936–5943, 2022, doi: 10.1109/access.2022.3142539. [5] y. zhao, w. li, f. liu, j. wang and a. munyole l., integrating heterogeneous structures and community semantics for unsupervised community detection in heterogeneous networks, expert systems with applications, https://doi.org/10.1016/j.eswa.2023.121821 [6] y. bai and j. wang, “observer-based distributed fault detection and isolation for second-order multi-agent systems using relative information,” j. franklin inst., vol. 358, no. 7, pp. 3779–3802, 2021, doi: 10.1016/j.jfranklin.2021.01.035. [7] p. c. gembarski and p. c. gembarski, “sciencedirect sciencedirect agent collaboration in a multi-agent-system for analysis and agent collaboration in a multi-agent-system for analysis and optimization of mechanical engineering parts optimization of mechanical engineering parts,” procedia comput. sci., vol. 176, pp. 592–601, 2020, doi: 10.1016/j.procs.2020.08.061. [8] j. cai , j. hao , h. yang , y. yang, x. zhao, y. xun and dongchao zhang , a new community detection method for simplified networks by combining structure and attribute information, expert systems with applications, https://doi.org/10.1016/j.eswa.2023.123103 [22] [9] r. budowle, e. krszjzaniek, and c. taylor, “students as change agents for community – university sustainability transition partnerships,” pp. 1–26, 2021. [10] a. belhadi, y. djenouri, g. srivastava, and j. c. lin, “reinforcement learning multi-agent system for faults diagnosis of mircoservices in industrial settings,” comput. commun., vol. 177, no. march, pp. 213–219, 2021, doi: 10.1016/j.comcom.2021.07.010. [11] d. grzonka, a. jakóbik, j. kołodziej, and s. pllana, “using a multi-agent system and artificial intelligence for monitoring and improving the cloud performance and security,” futur. gener. comput. syst., 2017, doi: 10.1016/j.future.2017.05.046. [12] d. liang, y. yang, r. li, and r. liu, “finite-frequency h − / h ∞ unknown input observer-based https://doi.org/10.1093/comnet/cnae001 https://www.sciencedirect.com/journal/expert-systems-with-applications https://www.sciencedirect.com/journal/expert-systems-with-applications https://doi.org/10.1016/j.eswa.2023.121821 https://www.sciencedirect.com/journal/expert-systems-with-applications https://doi.org/10.1016/j.eswa.2023.123103 vol. 6, no.1, january 2025 | 8 distributed fault detection for multi-agent systems,” j. franklin inst., vol. 358, no. 6, pp. 3258– 3275, 2021, doi: 10.1016/j.jfranklin.2021.01.042. [13] i. f. g. reis, i. gonçalves, m. a. r. lopes, and c. h. antunes, “a multi-agent system approach to exploit demand-side flexibility in an energy community,” util. policy, vol. 67, no. august, 2020, doi: 10.1016/j.jup.2020.101114. [14] z. ma, m. j. schultz, and m. værbak, “the application of ontologies in multi-agent systems in the energy sector : a scoping review,” pp. 1–31, 2019. [15] m. belaoued, a. derhab, a. khan, and s. mazouzi, “macomal : a multi-agent based collaborative mechanism for anti-malware assistance,” pp. 14329–14343, 2020. [16] v. s. de jesus, c. eduardo, f. manoel, j. viterbo, and e. bezerra, “bio-inspired protocols for embodied multi-agent systems,” vol. 1, no. icaart, pp. 312–320, 2021, doi: 10.5220/0010257803120320. [17] philippe gagnon, parallel algorithms for community detection in complex networks, hec montreal, 2017. [18] y. bai and j. wang, “fault detection and isolation using relative information for multi-agent systems,” isa trans., no. xxxx, 2021, doi: 10.1016/j.isatra.2021.01.030. [19] faiya et al., “a self organizing multi agent system for distributed voltage regulation,” vol. 3053, no. c, pp. 1–11, 2021, doi: 10.1109/tsg.2021.3070783. [20] m. a. shaik, c. science, and a. p. j. a. kalam, “agent-mb-divclues : multi agent mean based divisive clustering,” vol. 20, no. 5, pp. 5597–5603, 2021, doi: 10.17051/ilkonline.2021.05.629. [21] d. calvaresi, a. dubovitskaya, j. p. calbimonte, k. taveter, and m. schumacher, multi-agent systems and blockchain : results from a systematic literature review, vol. 2. springer international publishing. doi: 10.1007/978-3-319-94580-4. [22] s. chen, z. xia, h. li, j. liu, and h. pei, “controllable containment control of multi-agent systems based on hierarchical clustering,” int. j. control, vol. 0, no. 0, pp. 1–18, 2019, doi: 10.1080/00207179.2019.1610909. [23] i. a. saeed, a. l. i. selamat, and m. f. rohani, “a systematic state-of-the-art analysis of multiagent intrusion detection,” vol. 8, 2020, doi: 10.1109/access.2020.3027463. [24] n. s. elmitwally et al., “personality detection using context based emotions in cognitive agents,” 2022, doi: 10.32604/cmc.2022.021104. [25] m. h. qasem, n. obeid, a. hudaib, m. a. almaiah, and a. al-zahrani, “multi-agent system combined with distributed data mining for mutual collaboration classification,” 2021, doi: 10.1109/access.2021.3074125. [26] c. sweeney, e. ennis, m. mulvenna, r. bond, and s. o. neill, “how machine learning classification accuracy changes in a happiness dataset with different demographic groups,” 2022. [27] k. hamacher and r. buchkremer, “the application of artificial intelligence to automate sensory assessments combining pretrained transformers with word embedding based on the online sensory marketing index,” pp. 1–17, 2022. [28] k. griparic, m. polic, and s. member, “consensus-based distributed connectivity control in multi-agent systems,” ieee trans. netw. sci. eng., vol. 9, no. 3, pp. 1264–1281, 2022, doi: 10.1109/tnse.2021.3139045. vol.6, no.2, july 2025 | 108 p-issn: 2715-2448 | e-issn: 2715-7199 vol.6 no.2 july 2025 buana information technology and computer sciences (bit and cs) decision support system for the most chosen and preferred smartphone using the moora method rianna muhammad rizqy syahal maulana1, ryan hidayat2 1,2 faculty of technology law and business, sugeng hartono university, sukoharjo, indonesia e-mail: syahalmaulana1919@gmail.com1, hryan3480@gmail.com2 received: 2025/01/06 | revised: 2025/07/04 | accepted: 2025/07/30 abstract the rapid development of the digital era in indonesia has posed difficulties for consumers in choosing the right smartphone for them. this research aims to develop a dss (decision support system) using the moora (multi-objective optimization on the basis of ratio analysis) method to determine the best smartphone according to consumer preferences, with criteria including smartphone price, design, usability flexibility, durability, performance, and camera quality. evaluation is carried out by calculating the weight value of each criterion and ranking the smartphone alternatives based on a questionnaire. the research results show that the moora method is able to determine the best smartphone according to the people of solo and help consumers choose the right smartphone for them. this research involves five of the most widely used smartphones in indonesia (apple, samsung, xiaomi, oppo, vivo, and infinix) using the moora (multi-objective optimization on the basis of ratio analysis) method based on the defined criteria. the results show that xiaomi smartphones rank first both in terms of quality and user quantity. the practical implications of this research are significant, providing consumers with a data-driven approach to make informed smartphone choices, thereby enhancing their purchasing satisfaction. keywords: smartphone, dss, moora, preference i. introduction smartphones have become an essential part of everyday life, not only as communication tools but also as multifunctional devices encompassing entertainment, work, and education. in indonesia, the smartphone market continues to grow rapidly, driven by the growth of the younger generation and the increasing use of technology. smartphones themselves are electronic devices that function like mobile phones but are equipped with additional capabilities such as running applications, internet access, multimedia players, and various other features typically found on computers. smartphones generally use operating systems such as android or ios, which allow users to install additional applications as needed. according to the kamus besar bahasa indonesia (kbbi), a smartphone is "a smart phone, a mobile phone that has various computer functions, such as accessing the internet, receiving and sending emails, and so on." in 2024, it is estimated that the number of smartphone users in indonesia will reach 194.26 million people. this number has increased by 2.23% compared to 2023, which had 190.03 million users [1]. however, as the smartphone market in indonesia grows, consumers find it difficult to determine the right smartphone for them. the difficulty in determining a suitable smartphone necessitates the creation of a decision support system (dss) to assist consumers in selecting their smartphones based on the feedback provided by consumers this year through questionnaires distributed to the citizens of solo city. recent advancements in smartphone technology, such as the integration of ai and improved camera systems, have significantly influenced consumer preferences. understanding these trends is crucial for developing an effective dss that aligns with the evolving market dynamics. this study aims to create a dss that can analyze and recommend smartphones to consumers using the moora method. mailto:syahalmaulana1919@gmail.com mailto:hryan3480@gmail.com vol.6, no.2, july 2025 | 109 it is hoped that this research will provide consumers with a clear picture of the current favorite smartphones, along with clear and easily understandable data-driven analysis. ii. method 1. decision support system (dss) according to experts, a decision support system (dss) is an interactive system that supports the decision-making process by using a combination of data, models, and user interfaces (ui) to evaluate various decision scenarios. dss is designed to address problems with unique characteristics, requiring a flexible, interactive system that can be tailored to the user's needs. dss is typically used by middle to upper-level managers to aid in making strategic and tactical decisions involving many variables and uncertainties. [2] the main benefit of dss is to speed up the decision-making process. dss allows decision-makers to obtain accurate and up-to-date information, enabling them to make better and faster decisions. dss also helps reduce subjectivity in decision-making by providing objective and verifiable data. 2. research methodology the research method is a scientific way to obtain valid data, with the aim of finding, developing, and proving certain knowledge so that it can be used to understand, solve, and anticipate problems [4]. this study employs a quantitative approach. the data collection technique uses google forms by distributing them to individuals aged 18-30 in the solo area. the questionnaire was distributed to 104 respondents with 6 alternatives and 6 criteria. the data collection tool was developed with a closed questionnaire, namely a set of lists of statements or questions with possible answers that have been provided, so respondents only choose one of five alternative answers [5]. 3. moora methods the multi-objective optimization on the basis of ratio analysis (moora) method was first introduced by brauers and zavadkas [6]. moora is a multi-objective system developed to optimize multiple conflicting attributes at the same time [7],[9]. one of the advantages of the moora method is its ability to address objectives for conflicting criteria, where criteria can either be beneficial (benefit) or non-beneficial (cost) [10]. this method generates a final score for each option, which is then ranked based on the alternatives with the highest value. the steps involved in completing this procedure are as follows [11],[12]: a. preparing the decision matrix (1) b. calculating the normalization matrix (2) c. calculating the preference values in this step, which is the core of the process, each attribute is multiplied by the criteria weights for each alternative, then the results of the advantage criteria are added and subtracted from the results of the disadvantage criteria using the following formula [14],[15],[13]: vol.6, no.2, july 2025 | 110 (3) iii. results and discussion 1. application of alternatives in determining smartphone recommendations to help the community in choosing the right smartphone using the moora method, it starts with determining the alternative samples used. table 1. alternative data 2. application of criteria the following is the assessment criteria data from the decision support system in determining the most favorite smartphone using the moora method table 2. criteria description 3. alternative weights and criteria the following table 3 contains data on the alternative values of each criterion. table 3. alternative and criteria alternatif c1 c2 c3 c4 c5 c6 a1 2 4 4 4 3 4 a2 3 4 4 3 3 5 a3 4 4 4 4 4 4 a4 4 3 3 4 3 4 a5 3 4 3 2 4 4 a6 4 4 3 4 4 4 4. decision matrix the data in table 3 is converted into a matrix consisting of columns and rows to facilitate calculations in the next step, as presented below. alternative brand a1 apple a2 samsung a3 xiaomi a4 oppo a5 vivo a6 invinix criteria description weight value (wj) type c1 price 2 cost c2 flexible 1 benefit c3 design 1,5 benefit c4 performance 3,0 benefit c5 durable 1,5 benefit c6 camera 1 benefit vol.6, no.2, july 2025 | 111 𝑋 = [ 2 4 4 4 3 4 3 4 4 3 3 5 4 4 4 4 4 4 4 3 3 4 3 4 3 4 3 2 4 4 4 4 3 4 4 4] (4) 5. matrix normalization normalization aims to unite each matrix element so that the matrix elements have uniform values [16],[17]. here is the matrix normalization based on data in the decision matrix using equation 3 above: 𝑋𝑖𝑗 = [ 0,2390 0,4240 0,4618 0,4558 0,3464 0,3903 0,3585 0,4240 0,4618 0,3419 0,3464 0,4879 0,4781 0,4240 0,4618 0,4558 0,4618 0,3903 0,4781 0,3180 0,3464 0,4558 0,3464 0,3903 0,3585 0,4240 0,3464 0,2279 0,4618 0,3903 0,4781 0,4240 0,3464 0,4558 0,4618 0,3903] (5) 6. optimizing attribute the optimization value for each alternative is determined by summing the product of the criteria weights and the maximum attribute values (benefit type) and subtracting the sum of the product of the criteria weights and the minimum attribute values (cost type). [18]-[20]. 𝑋𝑖𝑗 = [ 0,2390(2) 0,4240(1) 0,4618(1,5) 0,4558(3) 0,3464(1,5) 0,3903(1) 0,3585(2) 0,4240(1) 0,4618(1,5) 0,3419(3) 0,3464(1,5) 0,4879(1) 0,4781(2) 0,4240(1) 0,4618(1,5) 0,4558(3) 0,4618(1,5) 0,3903(1) 0,4781(2) 0,3180(1) 0,3464(1,5) 0,4558(3) 0,3464(1,5) 0,3903(1) 0,3585(2) 0,4240(1) 0,3464(1,5) 0,2279(3) 0,4618(1,5) 0,3903(1) 0,4781(2) 0,4240(1) 0,3464(1,5) 0,4558(3) 0,4618(1,5) 0,3903(1)] (6) weighted matrix normalization results: 𝑋𝑖𝑗 = [ 0,4730 0,4240 0,6927 1,3674 0,5196 0,3903 0,7170 0,4240 0,6927 1,0257 0,5196 0,4879 0,9560 0,4240 0,6927 1,3674 0,6927 0,3903 0,9560 0,3180 0,5196 1,3674 0,5196 0,3903 0,7170 0,4240 0,5196 0,6837 0,6927 0,3903 0,9560 0,4240 0,5196 1,3674 0,6927 0,3903] (7) 7. rangking y value calculating the preference value for each alternative (student) using the moora formula, involves data normalization, determining the criterion weight, changing the criterion value into a matrix, and the preferences of each criterion [21]. the last stage in the dss process using the moora method is to determine the ranking [22]. tabel 4. nilai yi alternatif max min yi ranking a1 3,867 0,424 3,443 4 a2 3.866 0,424 3,442 5 a3 4,523 0,424 4,099 1 a4 4,070 0,318 3,752 3 a5 3,427 0,424 3,003 6 a6 4,350 0,424 3,926 2 vol.6, no.2, july 2025 | 112 the last stage in the dss process using the moora method is to determine the ranking [24]. tabel 5. rankingan alternatif merk mobil ranking xiaomi 1 infinix 2 oppo 3 apple 4 samsung 5 vivo 6 based on the calculation results using the moora method related to the smartphone brand, the alternative with code a1, the xiaomi handphone brand, is ranked 1st with a value of 4,099. iv. conclusion this study has succeeded in developing a decision support system (dss) with the moora method to identify the most favorite smartphone brands based on the preferences of smartphone users. the moora method excels in processing multiple criteria, simplifying data normalization, and producing objective final scores. references [1] muslim, a. (2024). pengguna smartphone ri diprediksi 194 juta. investor.id [2] adminbpmid. (2004). sistem pendukung keputusan (decision support system – dss). bpmid.uma.ac.id [3] eka, m. (2023). simak pengertian dan manfaat dari decision support systems (dss). it.telkomuniversity.ac.id [4] syaputra, a. e., & eirlangga, y. s. (2022). prediksi tingkat kunjungan pasien dengan menggunakan metode monte carlo. jurnal informasi dan teknologi, 97-102. [5] sari, m., rachman, h., astuti, n. j., afgani, m. w., & siroj, r. a. (2023). explanatory survey dalam metode penelitian deskriptif kuantitatif. jurnal pendidikan sains dan komputer, 3(01), 1016. [6] munthe, k., syahputra, t. r. a., pasuli, a. a., & hasibuan, m. a. (2022). sistem pendukung keputusan pemilihan pegawai honorer kelurahan medan sinembah menerapkan metode roc dan moora. bulletin of informatics and data science, 1(1), 20-29 [7] lubis, a. s., erwansyah, k., & rahmadiansyah, d. (2023). penerapan metode moora dalam menentukan kelayakan berita yang dapat ditayangkan di tvri sumatera utara. jurnal sistem informasi triguna dharma (jursi tgd), 2(3), 470-481. [8] zahara, a., & fakhriza, m. (2022). perbandingan metode smart, saw, moora pada pembangunan sistem pendukung keputusan pemilihan calon mitra statistik. journal of computers and digital business, 1(2), 72-82. [9] andri, r. h., & sitanggang, d. p. (2023). sistem penunjang keputusan (spk) pemilihan supplier terbaik dengan metode moora. jurnal sains informatika terapan, 2(3), 79-84. [10] rosita, i., & apriani, d. (2020). penerapan metode moora pada sistem pendukung keputusan pemilihan media promosi sekolah (studi kasus: smk airlangga balikpapan). metik jurnal, 4(2), 55-61. [11] yesinthia, v., siswanto, s., & kanedi, i. (2022). penerapan metode moora dalam penilaian kinerja guru di smk negeri 3 kota bengkulu. jurnal multidisiplin dehasen (mude), 1(1), 1319. [12] alatas, a., mumpuni, r., & nurlaili, a. l. (2021). spk penilaian kinerja untuk kenaikan jabatan pegawai menggunakan metode moora. jurnal informatika dan sistem informasi, 2(2), 171-180. [13] aritonang, h. d., azmi, z., & suherdi, d. (2024). penerapan metode moora dalam menentukan kelayakan barang return to vendor (rtv) kepada distributor. jurnal sistem informasi triguna dharma (jursi tgd), 3(6), 1051-1062. vol.6, no.2, july 2025 | 113 [14] kepuasan konsumen sebagai variabel intervening pada pengguna transportasi migo di surabaya. jurnal pendidikan tata niaga (jptn), 8(2). [15] syaputra, a. e., & eirlangga, y. s. (2022). prediksi tingkat kunjungan pasien dengan menggunakan metode monte carlo. jurnal informasi dan teknologi, 97-102. [16] rohman, a. a., bachri, o. s., & wahyuningsih, p. (2024). penentuan karyawan terbaik menggunakan metode moora pada swalayan m di kota tegal. jurnal ilmiah intech: information technology journal of umus, 6(1), 36-41. [17] nasyuha, a. h., ramadhan, m., & syahril, m. (2021). kelayakan event organizer pada event ibc (intensive bible course) di yayasan giving indonesia menggunakan metode moora. device: journal of information system, computer science and information technology, 2(1), 1-8. [18] tarigan, n. m. b., sinaga, b., amelia, r., & krisswanti, y. (2024). penerapan metode moora dalam pemilihan bibit lele terbaik. jurnal media informatika, 6(1), 161-169. [19] situmorang, d. m. s., nugroho, n. b., winata, h., erwansyah, k., & santoso, i. (2024). penerapan sistem pendukung keputusan penerimaan taruna/i menggunakan metode moora. jurnal teknologi sistem informasi dan sistem komputer tgd, 7(2), 207-218. [20] arshad, m. w. (2024). implementation of entropy and additive ratio assessment methods in determining the best warehouse location. bulletin of computer science research, 4(4), 318326. [21] romlah, s. r., lutfi, a., & lidimillah, l. f. (2024). implementasi metode moora dalam sistem pendukung keputusan pemilihan siswa terbaik di mi at-taqwa bondowoso. prosiding seminastika, 5(1), 96-100. [22] amanda, a., muhazir, a., & kustini, r. (2024). penerapan metode moora dalam pemilihan anggota staff redaksi. jurnal sistem informasi triguna dharma (jursi tgd), 3(4), 583-591 vol. 6, no.2, july 2025 | 56 p-issn: 2715-2448 | e-issn: 2715-7199 vol.6 no.2 july 2025 buana information technology and computer sciences (bit and cs) development of a deep learning-based text-to-speech system for the malang walikan language using the pre-trained speecht5 and hifi-gan models aina avrilia imani1, aviv yuniar rahman2, firman nurdiyansyah3 1 2 3 department of informatic engineering, universitas widya gama, malang, indonesia e-mail: avriliaaina@gmail.com1, aviv@widyagama.ac.id2, firmannurdiyansyah@widyagama.ac.id3 received: 2025/06/07 | revised: 2025/06/27 | accepted: 2025/07/27 abstract the walikan language of malang is a form of local cultural heritage that needs to be preserved in the digital era. this study aims to develop and evaluate a deep learning-based text-to-speech (tts) system capable of generating speech in the walikan language of malang using pre-trained speecht5 and hifigan models without fine-tuning. in this system, speecht5 is used to convert text into mel-spectrograms, while hifi-gan acts as a vocoder to generate audio signals from the mel-spectrograms. the dataset used consists of 1,000 sentences in the walikan language of malang. the system evaluation was carried out using objective metrics of word error rate (wer) and character error rate (cer), by comparing the results of synthetic audio transcriptions against two types of reference audio, namely the original voices of female speakers and male speakers, using the automatic speech recognition (asr) system. the female voice was recorded with controlled articulation, while the male voice used natural intonation in everyday conversation. the results show that synthetic audio has the highest error rate with a wer of 0.9786 and a cer of 0.9024. meanwhile, female audio has a wer of 0.5471 and a cer of 0.1822, while male audio shows a wer of 0.6311 and a cer of 0.2541. these findings indicate that the tts model without fine-tuning is not yet capable of producing synthetic voices that can be recognized accurately by the asr system, especially for regional languages that are not included in the initial training data. therefore, the fine-tuning process and the preparation of a more representative dataset are important so that the tts system can support the preservation of the walikan malang language more effectively in the digital era. keywords: hifi-gan, malang walikan language, speecht5, text-to-speech, and word error rate. i. introduction the walikan language of malang is a local cultural heritage that deserves preservation, especially in today's digital era, where regional languages are increasingly threatened with extinction. deep learning-based text-to-speech (tts) technology has developed rapidly and is one solution for language preservation by converting text into natural speech. however, most existing tts models focus on resource-intensive languages, making them less than optimal for low-resource regional languages like walikan. research on tts for regional languages is still very limited, especially in cases where there is little or no training data (zero-shot). walikan also has unique characteristics not found in other languages, making it crucial to test how pre-trained models like speecht5 and hifi-gan perform under these conditions without fine-tuning. several previous studies have used speecht5 and hifi-gan models for languages with more resources and have shown quite good results [1], [2]. however, the use of pre-trained models directly (zero-shot) in low-resource regional languages, particularly malang walikan, is still rarely researched. vol. 6, no.2, july 2025 | 57 other studies have compared objective evaluation methods such as word error rate (wer) with subjective mean opinion score (mos) assessments to assess synthetic speech quality [3]. this study aims to develop and evaluate a deep learning-based tts system using pre-trained models speecht5 and hifi-gan for malang walikan without fine-tuning. the evaluation was conducted using objective metrics wer and character error rate (cer) by comparing the synthetic results to real human speech to assess the model's performance on low-resource languages. in the development of deep learning-based text-to-speech systems, various architectures have been developed to produce voices that increasingly resemble human speech. one recent model that has demonstrated superior performance is speecht5 [4], a transformer-based pre-trained model with an encoder-decoder approach designed for various speech processing tasks, including speech synthesis. this model is capable of understanding text structure and converting it into a mel-spectrogram representation, taking into account intonation and rhythm. to generate the final speech signal, hifigan [5] was used, a generative adversarial network (gan)-based vocoder designed to convert melspectrograms into high-quality, natural audio. the combination of these two models was found to be effective in generating synthetic speech that approximates human voice quality. this research provides an important contribution to understanding how large pre-trained models perform on regional languages not yet represented in the training data, and provides a basis for further development through fine-tuning and the collection of a more comprehensive dataset, to support the digital preservation of the walikan malang language. ii. methods the steps in this research are shown in (figure 1). . figure 1. steps a research. (source: personal preparation) the dataset used in this study consists of 1,000 sentences in the walikan language of malang. each sentence was recorded by two different speakers, one male and one female, resulting in a total of 2,000 audio files. this dataset was used to study the linguistic patterns and acoustic characteristics of walikan and to evaluate the system's performance on voice variations based on speaker gender. the collected data was stored in spreadsheet format (excel) for ease of management and processing. before use, the data underwent pre-processing to remove irrelevant symbols or emojis and normalize the text. an example of the data used is shown in table 1. evaluation vol. 6, no.2, july 2025 | 58 table 1. dataset language walikan language walikan language indonesian etas maya sing cedek gang iku asaib sam rasane sate atam yang dekat gang itu biasa mas rasanya yang ayahab kuwi gluduk arudam yang bahaya itu petir madura tambah nade ae nawak kotis iki tambah gila aja teman satu ini wah umak lihai juga berbahasa arudam wah kamu lihai juga berbahasa madura rintep pol kakakku pintar banget kakakku agomes warga licek sukses halokes e semoga warga kecil sukses sekolah nya halokes sek jum sekolah dulu jum nde wendit akeh sedeb sam di wendit banyak monyet mas likis e adikku mari dicokot kucing kaki nya adikku habis digigit kucing bangga nggae soak arema bangga pakai kaos arema 1. pre-trained speecht5 model speecht5 is a pre-trained model based on a transformer encoder-decoder architecture used to convert text into a mel-spectrogram, a visual representation of audio frequencies over time. this model was implemented without fine-tuning, using the default configuration from the hugging face library. the speecht5 model uses a transformer-based encoder-decoder architecture, with six additional prenet and post-net modules to process input and output in the form of text and speech. the encoder consists of two parts: the speech encoder pre-net to convert the raw speech signal into speech features, and the text encoder pre-net to convert the text into a textual representation. the outputs of these two parts are combined into a unified representation that is used by the decoder. figure 2. speecht5 architecture the decoder also has two paths: the speech decoder, which generates a mel-spectrogram from text input, and the text decoder, which generates text from voice input. the pre-net and post-net components in each section enhance the quality of the representation and output. the speech encoder pre-net utilizes features from wav2vec 2.0, while the decoder utilizes speaker embeddings (x-vectors) to support multi-speaker synthesis. the initial pre-training process is multi-task and cross-modal, including bidirectional masking prediction, sequence-to-sequence generation, and text pre-training using an infilling strategy. the encoder output representation is discretized using vector quantization, allowing the model to align the voice and text modalities. vol. 6, no.2, july 2025 | 59 the implementation of the speecht5 model as the core of the text-to-speech system for the malang walikan language is carried out using the following steps: a. data preparation and pre-trained model the pre-trained speecht5 model is used in the walikan language sentence dataset without finetuning. the audio output is evaluated using asr and word error rate (wer). b. text preprocessing 1) normalization: standardize the text (lowercase, remove irrelevant symbols/punctuation). 2) tokenization: break the text into tokens (words/subwords). 3) conversion to numeric id: map tokens to unique numbers based on the vocabulary. 4) embedding: convert the id tokens into a numeric representation vector that the model understands. c. encoding with a transformer encoder the embedding vector is processed by a transformer encoder with self-attention to produce a contextual representation of the text (hidden states). d. mel-spectrogram prediction by the decoder the transformer decoder converts the contextual representation into a mel-spectrogram through self-attention and cross-attention with the encoder output. the mel-spectrogram depicts sound energy at various frequencies over time. e. mel-spectrogram conversion to audio the decoded mel-spectrogram is converted to audio using hifi-gan. 2. hifi-gan hifi-gan is used as a vocoder to convert the mel-spectrogram output from speecht5 into a highquality audio signal. this model plays a role in the final stage of the synthesis process, producing natural-sounding voices that approximate human voice quality. the output, in the form of a melspectrogram from speecht5, serves as input to hifi-gan, which then produces the final audio in wav format. hifi-gan [6] is a generative adversarial network (gan)-based vocoder designed to convert mel-spectrograms into high-quality raw audio signals. its architecture consists of a generator and two types of discriminators: the multi-scale discriminator (msd) and the multi-period discriminator (mpd) [7]. training is conducted in an adversarial manner, where the generator and discriminator compete to produce increasingly natural sounds. figure 3. hifi-gan architecture [5] the hifi-gan generator is a fully convolutional network with an upsampling process based on transposed convolution [8]. to capture complex temporal patterns, a multi-receptive field fusion (mrf) module is used, which combines multiple residual blocks with different kernels and dilations. vol. 6, no.2, july 2025 | 60 mpd focuses on periodic patterns in audio signals with sub-discriminators that process the input based on specific time periods (e.g., 2, 3, 5, 7, 11). meanwhile, msd captures patterns at multiple temporal scales, processing audio at different resolutions (raw, average pooled ×2, and ×4) to identify global characteristics of the signal. 3. system output the system generates synthetic audio in malang walikan, reflecting the phonetic and linguistic characteristics of the input text. the resulting audio can be listened to as a spoken representation of the sentences in the dataset. 4. evaluation the evaluation was conducted using the word error rate (wer) and character error rate (cer) metrics, which measure the error rate between the automatic transcription of synthetic audio and the original text. the lower the wer and cer values, the higher the system's accuracy and naturalness in producing speech that matches the text. evaluation of text-to-speech (tts) results is performed by measuring the accuracy of the synthesized audio transcription against the reference text using two main metrics: the word error rate (wer) and the character error rate (cer). a. word error rate the word error rate (wer) is the primary metric for evaluating the accuracy of a text-to-speech (tts) system by comparing the audio transcription output from the model with the original text. the wer is calculated based on the number of word substitutions (s), deletions (d), and insertions (i) divided by the total number of words (n) in the original text, expressed as a percentage: wer = 𝑆+𝐷+𝐼 𝑁 𝑥 100 if wer = 0%, the system produces perfect output that is identical to the original text. the higher the wer value, the greater the errors in the speech synthesis results [9], [10]. b. character error rate character error rate (cer) is an evaluation metric similar to wer, but calculated at the character level. cer is more sensitive to small errors, especially in short sentences or languages with nonstandard structures, such as walikan. the cer formula is similar to wer, namely: cer = 𝑆+𝐷+𝐼 𝑁 𝑥 100 where s, d, and i are the number of character substitutions, deletions, and insertions, and n is the total number of characters in the original text. cer complements wer by providing a more detailed error analysis at the phonetic level [11] iii. results and discussions 1. word error rate synthetic audio quality was evaluated by calculating the word error rate (wer) for nine representative test sentences with varying error levels. three types of audios were analyzed: the speecht5 + hifi-gan synthesized audio, a female voice, and a male voice, to determine how well the asr system recognized each type of voice. (1) (2) vol. 6, no.2, july 2025 | 61 table 2. word error rate. test sentence wer speecht5 + hifi-gan wer voice female wer voice male info lokasi dong nawak 0.5 0.25 0.25 sesuai dengan gambar di bis halokes 0.5 0.17 0.33 wah umak lihai juga berbahasa arudam 0.67 0.5 0.83 rame ilakes area kayutangan di malam minggu 0.71 0.29 0.71 sekali nade tetep nade 0.75 0.75 0.75 agomes lancar rejeki hari ini 0.8 0.6 0.4 nakam lah mbah 1.0 0.33 1.0 umak ngalam 1.0 0.5 1.0 jenenge ae ongis nade yo nade temenan 1.0 0.57 0.71 table 2 shows examples of wer values for the three audio types. the results indicate that all synthetic audio had wer values above 0.30, indicating a high level of pronunciation errors. four sentences achieved a maximum wer of 1.0, meaning that all words were not recognized by asr. the sentences with the lowest wer were "info lokasi dong nawak" and "sesuai dengan gambar di bis halokes," each with a value of 0.5. in contrast, human voices performed better and more consistently, with some sentences achieving wers as low as 0.25. this indicates that natural audio is more capable of conveying phonetic information accurately. the high error rate in synthetic audio is likely due to the striking phonetic differences and the model's limitations in handling the non-standard walikan language structure and its inversions. sentences such as "wah umak lihai juga bahasa arudam" reflect linguistic challenges for the model that has not undergone fine-tuning. as a follow-up to this evaluation, the scatter plot in figure 4 was used to visualize the distribution of wer values for each test sentence by voice type category. this visualization provides a more comprehensive picture of the distribution of recognition errors at the sentence level. figure 4. scatter plot comparing wer values by audio type figure 4 shows a comparison of the word error rate (wer) values for nine test sentences based on three audio types: the ones synthesized by the speecht5 + hifi-gan model (red/orange), the vol. 6, no.2, july 2025 | 62 female voice (blue), and the male voice (green). in general, the synthetic audio produced the highest and most consistent wer values, with all sentences having values above 0.5, even reaching 1.0 in the last three sentences. this indicates that the text-to-speech system is not yet capable of producing pronunciations close to human accuracy, especially in the complex context of walikan. meanwhile, human audio, particularly female voices, demonstrated better and more stable performance, with wer values ranging from 0.2 to 0.6. male voices also tended to be more accurate than synthetic audio, but experienced greater fluctuations between sentences. this difference indicates that synthetic audio still faces significant challenges in matching the natural phonetic characteristics of human speakers. figure 5. histogram of wer comparison figure 5 shows a comparison of the distribution of word error rate (wer) values for three types of audios: female, male, and the speecht5 + hifi-gan synthesized voices. the histogram shows that the wer values of the synthesized voices are mostly close to 1.0, indicating asr's failure to recognize the sentence content. in contrast, human voices show a lower and more even distribution of wer values, some even approaching 0.0, indicating near-perfect transcription. this finding confirms that without fine-tuning, the speecht5 + hifi-gan model is not yet capable of producing pronunciations accurate enough for asr to recognize. table 3. average word error rate average word error rate (wer) speecht5 + hifigan real voice (female) real (male) 0.9786 0.5471 0.6311 testing 1,000 test sentences showed an average wer of 0.9786 for the synthetic audio from speecht5 + hifi-gan (table 3), close to 1.0, indicating that the majority of words were not recognized by asr. this reflects the low quality of the synthesis, especially in the context of walikan, which has unusual phonetics. in contrast, human voices performed better, with an average wer of 0.5471 (female) and 0.6311 (male), a difference of ±0.33–0.43 lower than the synthetic results. this confirms that asr is better able to recognize natural voices. vol. 6, no.2, july 2025 | 63 in addition to inaccurate pronunciation, the synthetic results also showed anomalies such as random capitalization and phonetic errors due to the model's bias toward english, for example, "sam" being recognized as "some". 2. character error rate evaluation using the character error rate (cer) provides a more detailed picture of phonetic errors at the character level, especially for minor errors such as letter substitutions or spelling mistakes. table 4 summarizes the minimum, maximum, median, and average cer values for the three audio source categories tested. table 4. character error rate evaluation voice category minimum cer maximum cer median cer mean cer speecht5 + hifi-gan 0.3243 1.0000 1.0000 0.9024 voice female 0.0000 1.6154 0.1461 0.1822 voice male 0.0000 1.0000 0.2143 0.2541 character error rate analysis (table 4) shows that the synthetic audio from the speecht5 + hifigan model performed poorly compared to human audio. the average cer was 0.9024, with a median and maximum of 1.00, indicating that many sentences failed to be recognized by asr. even the minimum value of 0.3243 still indicated significant errors at the character level. in contrast, human audio was much more accurate. female voices had an average cer of 0.1822 and male voices 0.2541, with a minimum value of 0.00 for both, indicating some sentences were perfectly recognized. these results align with previous wer findings, indicating that errors in synthetic audio occur not only at the word level but also at the character level. this reflects the model's limitations in representing walikan phonology without fine-tuning. based on the evaluation, the speecht5 and hifi-gan-based tts systems performed poorly when used on walikan without fine-tuning. high wer and cer values for synthetic audio indicate that asr has difficulty recognizing the model's output, in contrast to human speech, which produces better accuracy. table 5. comparison of regional language tts studies author dataset data ground truth method evaluation validate wer cer mos agustina, c., 2024. [12] 250 sentence language banjar vits 3.604 alhuda, m.y., 2025. [13] 500 sentence language palembang vits 4.58 our proposed 1000 sentence language walikan malang voice artificial speecht5 & hifi-gan 0,9786 0,9024 voice female speecht5 & hifi-gan 0,5471 0,1822 voice male speecht5 & hifi-gan 0,6311 0,2541 vol. 6, no.2, july 2025 | 64 table 5 compares this research with other tts studies on regional languages. the studies by agustina (2024) and alhuda (2025) used vits with a smaller dataset and only assessed voice quality through mos (3.604 and 4.58), without objective metrics like wer or cer. meanwhile, this study used 1,000 walikan sentences and objectively evaluated them using wer and cer. the results showed that the synthetic audio had a wer of 0.9786 and a cer of 0.9024, indicating significant errors. however, the female and male voices recorded wers of 0.5471 and 0.6311, and cers of 0.1822 and 0.2541, indicating better asr recognition of human voices. the advantage of this approach lies in the use of more measurable objective metrics and the hybrid speecht5 + hifi-gan model, which has potential, although it still requires further optimization and training to handle the complexity of regional languages and natural voice variations. iv. conclusions this study developed a preliminary tts system for malang walikan using speecht5 and hifigan without fine-tuning. although the synthesis results were not optimal, the system demonstrated basic potential for improvement through further training with local data. evaluation using asr-based wer and cer showed that the synthetic audio (speecht5 + hifi-gan) produced a wer of 0.9786 and a cer of 0.9024, indicating a very high error rate. audio from a female voice had a wer of 0.5471 and a cer of 0.1822, while audio from a male voice recorded a wer of 0.631 and a cer of 0.2541. these differences indicate that voice characteristics significantly impact system performance. the system was unable to replicate the speaker's voice characteristics, such as intonation and clarity of pronunciation, likely due to the lack of fine-tuning and a lack of adaptation to the unique structure of walikan. however, the pronunciation of some words was quite accurate, indicating potential for further development.this research is an initial contribution to the development of tts for undocumented regional languages, and emphasizes the importance of adapting models to local contexts for more accurate and natural results. vol. 6, no.2, july 2025 | 65 references [1] j. lehečka, z. hanzlíček, j. matoušek, and d. tihelka, “zero-shot vs. few-shot multi-speaker tts using pre-trained czech speecht5 model,” pp. 46–57, 2024, doi: 10.1007/978-3-03170566-3_5. [2] h. wang, “understanding zero-shot rare word recognition improvements through llm integration,” 2025, [online]. available: http://arxiv.org/abs/2502.16142 [3] a. kirkland, s. mehta, h. lameris, g. e. henter, e. szekely, and j. gustafson, “stuck in the mos pit: a critical analysis of mos test methodology in tts evaluation,” no. august, pp. 41– 47, 2023, doi: 10.21437/ssw.2023-7. [4] j. ao et al., “speecht5: unified-modal encoder-decoder pre-training for spoken language processing,” oct. 2021, [online]. available: http://arxiv.org/abs/2110.07205 [5] j. kong, j. kim, and j. bae, “hifi-gan: generative adversarial networks for efficient and high fidelity speech synthesis,” oct. 2020, [online]. available: http://arxiv.org/abs/2010.05646 [6] j. su, z. jin, and a. finkelstein, “hifi-gan: high-fidelity denoising and dereverberation based on speech deep features in adversarial networks,” jun. 2020, [online]. available: http://arxiv.org/abs/2006.05694 [7] z. qiu, j. tang, y. zhang, j. li, and x. bai, “a voice cloning method based on the improved hifi-gan model,” comput intell neurosci, vol. 2022, 2022, doi: 10.1155/2022/6707304. [8] d. lim, s. jung, and e. kim, “jets: jointly training fastspeech2 and hifi-gan for end to end text to speech,” proceedings of the annual conference of the international speech communication association, interspeech, vol. 2022-septe, pp. 21–25, 2022, doi: 10.21437/interspeech.2022-10294. [9] a. ali and s. renals, “word error rate estimation without asr output: e-wer2,” proceedings of the annual conference of the international speech communication association, interspeech, vol. 2020-octob, pp. 616–620, 2020, doi: 10.21437/interspeech.2020-2357. [10] a. ali and s. renals, “word error rate estimation for speech recognition: e-wer,” jul. 2018. [online]. available: https://github.com/qcri/e-wer [11] i. kottayam and j. james, “advocating character error rate for multilingual asr evaluation,” 2023. [12] c. agustina, “implementasi teknologi text to speech bahasa banjar menggunakan metode vits,” (doctoral dissertation, universitas islam negeri sultan syarif kasim riau)., 2024. [13] m. y. alhuda, “text to speech bahasa palembang menggunakan metode vits,” (doctoral dissertation, universitas islam negeri sultan syarif kasim riau)., 2025. vol.6, no.2, july 2025 | 66 p-issn: 2715-2448 | e-issn: 2715-7199 vol.6 no.2 july 2025 buana information technology and computer sciences (bit and cs) classification of rice plant diseases based on leaf images using the multi class support vector machine (m-svm) method febiana angela tanesab 1, rangga pahlevi putra 2, aviv yuniar rahman 3 1 2 3 department of informatic engineering, universitas widya gama, malang, indonesia e-mail: email: febytanesab@gmail.ac.id 1, rangga@widyagama.ac.id 2 , aviv@widyagama.ac.id 3 received: 2025/05/17 | revised: 2025/06/26 | accepted: 2025/07/27 abstract the rice farming sector plays an important role in the indonesian economy, considering that rice is the main staple food. according to irri, rice farmers experience crop losses of up to 37% each year due to pests and diseases. this study aims to classify rice plant diseases using the multi-class support vector machine (m-svm) method based on leaf images. this study aims to provide education to farmers in recognizing and overcoming diseases in rice plant leaves. the types of rice leaf diseases classified in this study include blast, kresek, and tungro. the data used in this study amounted to 1200, which were divided by varying training and testing data ratios, from 10% training and 90% testing to 90% training and 10% testing. each variation of features and data division was evaluated by calculating the model performance parameters. the features used for classification include color (rgb) and texture (glcm) from leaf images. the test results showed that the best accuracy obtained was 85.5% using a combination of color and texture features. keywords: accuracy, disease classification, glcm, leaf image, m-svm, rice. i. introduction the rice farming sector plays an important role in contributing to the indonesian economy, because rice is one of the largest commodities. many countries, including indonesia, make rice their main staple food. therefore, indonesia needs to continue to innovate so that the rice supply remains abundant and stable [1]. agriculture itself is an activity that utilizes nature to produce food, one of which is rice cultivation. however, rice plants are often attacked by various diseases, such as leaf blight (kresek), blast, tungro and others [2]. the development of digital image processing technology and artificial intelligence (ai) provides potential solutions in the agricultural sector, especially in terms of identifying plant diseases. with the help of machine learning algorithms, such as support vector machine (svm), the classification process can be carried out [3], multi-class support vector machine (msvm) is a variant of the support vector machine (svm) method used to solve multi-class classification problems. svm is basically a classification algorithm designed to handle two-class problems (binary classification) [4] based on the background of this problem, researchers propose a solution by using the multi-class support vector machine (m-svm) method. the use of the m-svm algorithm allows disease classification based on patterns and textures on rice leaves, so that each type of disease can be recognized more quickly and accurately. ii. methods in (figure 1) it will explain the research stages including several steps carried out systematically to achieve the objectives of the research, the research stages include starting, input of rice leaves, analysis of the problem identification system, implementation, trial, success. mailto:febytanesab@gmail.ac.id mailto:rangga@widyagama.ac.id mailto:aviv@widyagama.ac.id vol.6, no.2, july 2025 | 67 . figure 1. research stage flowcart. (source: personal preparation) 1. input data for rice disease leaves: (a). leaf blight (kresek) (b). blast (c). tungro figure 2. image of rice leaves in (figure 2) we will explain about 3 diseases of rice as follows: a. bacterial leaf blight is a very common disease found in rice fields. the main cause of this disease is the bacteria xanthomonas oryzae. symptoms of bacterial leaf blight on leaf blades are characterized by damage that usually begins a few centimeters from the edge, which appears as lines and blisters, then spreads to the wavy edges [ 1]. b. blast disease caused by pyricularia grisea is an important disease in rice plants in indonesia, especially in upland rice in dry land. grisea infects the leaves and causes disease symptoms in the form of diamond-shaped brown spots called leaf blast [5]. c. tungro is a disease caused by a double infection of 2 different types of viruses. the second virus in question is rice tungro spherical virus (rtsv) and rice tungro bacilliform virus (rtbv). vol.6, no.2, july 2025 | 68 symptoms of tungro disease are that the leaves will turn yellow starting from the tips of the leaves that are still in the growth stage [5]. rice plants are susceptible to various types of diseases. in this study, we focus on three main types of diseases in rice plants, namely tungro disease, leaf blight, and leaf blast. the data used consists of 1200 leaf images divided into 3 classes, namely 400 leaf blight (kresek) image data, 400 leaf blast image data, 400 tungro image data. data division is carried out for training data (80%) and test data (20%) [3]. 2. pre-processing pre-processing is a crucial step that is carried out before the image is used for feature extraction or classification model development. the main purpose of this stage is to prepare the image so that it is more ready for further analysis and can improve accuracy [6]. the disease detection system at this processing stage includes: normalization, contraction, cropping, resize. 3. glcm and rgb feature extraction glcm (gray level co-occurrence matrix) and rgb (red, green, blue) feature extraction are used to analyse the texture and color of rice leaf images, especially in detecting and classifying diseases that attack rice leaves [7]. when rice leaves are infected with disease, both the texture and color of the leaves will experience different changes from healthy leaves. infection can cause color changes, such as yellowish or brownish, as well as the appearance of spots with certain intensities, which can be analyzed through rgb features [8]. in addition, changes in texture patterns such as spots, lines, or holes on leaves can be evaluated using the glcm method which extracts features such as contrast and correlation. by combining texture analysis using glcm and color analysis using rgb, we can obtain more complete information about the condition of rice leaves, thereby increasing the accuracy of disease identification and classification [9]. 4. m-svm classification classification using the multi-class svm (msvm) method is divided into two stages, namely training and testing, where the image dataset goes through a feature extraction process using the glcm and rgb methods. glcm is used to extract texture information, such as contrast and correlation, while rgb is used to analyse colour characteristics in rice leaf images. furthermore, the extracted images are classified using multi-class svm. through this classification process, it can be identified whether the input leaves are included in the category of normal leaves or diseased rice leaves, and can be separated based on their respective classes by considering a combination of texture and colour features to improve identification accuracy [10],[11],[12]. 5. accuracy evaluation at this stage there are several steps taken, namely: a. confusion matrix use a confusion matrix to see how well the model classifies diseases. the confusion matrix will show the number of correct and incorrect predictions for each disease class (e.g., leaf blight, leaf blast, and tungro) that exist [9], [13], [14]. b. accuracy accuracy = 𝑇𝑃+𝑇𝑁 𝑇𝑃+𝑇𝑁+𝐹𝑃+𝐹𝑁 × 100% (1) vol.6, no.2, july 2025 | 69 c. precision precision = 𝑇𝑃 𝑇𝑃+𝐹𝑃 × 100% (2) d. recall recall = 𝑇𝑃 𝑇𝑃+𝐹𝑁 × 100% (3) the research will be conducted at jl. sudimoro, behind the sawah cafe, and is targeted to take place from november 2024 to february 2025. this extended period will provide sufficient time for the researcher to make necessary preparations, such as understanding the research problems, objectives, methods, and the tools required for the research process [15]. iii. results and discussions a. feature extraction using gray level co-occurrence matrix (glcm) and rgb feature extraction using the gray level co-occurrence matrix (glcm) method is a technique in image processing used to obtain texture information from images. glcm analyzes the spatial relationship between pixels based on gray levels to form a co-occurrence matrix, from which texture features such as contrast, correlation can be calculated. at this stage, the method is used to detect rice leaf images. in addition, the basic color values of the image are also used by taking rgb (red, green, blue) values directly from each pixel as additional features that represent image color information. figure 3. blast image capture interface design in (figure 3.) to convert data from image form into numeric form, a data conversion process design is needed that utilizes a graphical interface (gui) using the matlab programming language. this process is important for the purposes of texture analysis with the m-svm method. the blast, kresek and tungro data values obtained from the feature extraction results consist of several fields that can be displayed in (table 1). tabel 1. blast, kresek and tungro image datasets project_id contrast correlation red greend blue results blast111.jpg 0.091126 0.96395 0.7276 0.74142 0.75599 blast blast144.jpg 0.09572 0.95864 0.74756 0.72766 0.69713 blast vol.6, no.2, july 2025 | 70 kresek11.jpg 0.15015 0.93488 0.56585 0.62566 0.76428 plastic bag kresek109.jpg 0.1193 0.95152 0.74032 0.72572 0.71702 plastic bag tungro1.jpg 0.073608 0.96713 0.7847 0.72117 0.66278 tungro tungro100.jpg 0.053928 0.97586 0.79111 0.73136 0.65979 tungro b. training data multiclass support vector machine (m-svm) method table 2. results of polynomial m-svm training evaluation m-svm(polynomial) split ratio accuracy precision recall data training testing train test 10% 90% 88% 83.3751% 83.3333% 120 1080 20% 80% 89.7222% 84.5861% 84.5833% 240 960 30% 70% 85.1852% 78.8981% 77.7778% 360 840 40% 60% 85.5556% 78.7160% 78.3333% 480 720 50% 50% 85.7778% 78.5025% 78.6667% 600 600 60% 40% 85.2778% 77.8825% 77.9167% 720 480 70% 30% 85.2381% 79.2707% 77.8571% 840 360 80% 20% 84.5833% 77.1485% 76.8750% 960 240 90% 10% 85.1852% 77.9073% 77.7778% 1080 120 in (table 2) it is explained that the results of the training data evaluation using the m-svm polynomial method obtained a high accuracy score, namely at a split ratio of 20:80, the number of training data is 240 with an accuracy score of 89.7222%. precision 84.5861% and recall 84.5833. and the results of the m-svm polynomial graph and the calculation of the confusion matrix with the highest value at a split ratio of 20:80, can be seen in (figure 4) and (figure 5) below. figure 4. m-svm polynomial 20:80 graph figure 5. cm m-svm polynomial 20:80 vol.6, no.2, july 2025 | 71 table 3. linear m-svm training evaluation results m-svm(linear) split ratio accuracy precision recall data training testing train test 10% 90% 87.2222% 81.2963% 80.8333% 120 1080 20% 80% 89.4444% 84.1667% 84.1667% 240 960 30% 70% 84.8148% 77.3130% 77.2222% 360 840 40% 60% 86.2500% 79.5291% 79.3750% 480 720 50% 50% 84.4444% 76.7066% 76.6667% 600 600 60% 40% 84.9074% 77.2398% 77.3611% 720 480 70% 30% 84.2063% 76.4328% 76.3095% 840 360 80% 20% 84.5833% 76.7444% 76.8750% 960 240 90% 10% 84.3827% 76.7453% 76.5741% 1080 120 in (table 3), it can be seen that the m-svm linear method shows performance variations at various split ratios of training and testing data. at a split ratio of 20:80, the m-svm linear model obtained very good results with the highest accuracy score of 89.4444%, followed by a precision score of 84.1667% and a recall of 84.1667%. and the results of the m-svm linear graph and the calculation of the confusion matrix with the highest value at a split ratio of 20:80, can be seen in (figure 6) and (figure 7) below. figure 6. m-svm linear graph 20:80 figure 7. cm m-svm linear 20:80 table 4. results of gausian m-svm performance evaluation m-svm(polynomial) split ratio accuracy precision recall data training testing train test 10% 90% 80% 77.5360% 76.6667% 120 1080 20% 80% 79.2125% 75.7534% 75.8333% 240 960 30% 70% 79% 69.3777% 70% 360 840 40% 60% 77.2222% 65.0255% 65.8333% 480 720 vol.6, no.2, july 2025 | 72 50% 50% 76% 62.9947% 64% 600 600 60% 40% 74.6296% 60.8490% 59.5833% 720 480 70% 30% 71.1905% 55.6062% 56.7857% 840 360 80% 20% 71.5278% 54.3445% 57.2917% 960 240 90% 10% 71.2346% 53.6462% 56.8519% 1080 120 in (table 4) it is explained that the gaussian m-svm method on training data obtained the highest accuracy at a split ratio of 10:90, with a total of 120 data. at this split ratio, the accuracy obtained was 80%, precision 75.7534%, and recall 75.8333%. the gaussian m-svm graph and confusion matrix calculation for a split ratio of 10:90 can be seen in (figure 8) and (figure 9). figure 8 gaussian m-svm graph 10:90 figure 9 gaussian cm m-svm 10:90 c. test data for the multiclass support vector machine (m-svm) method table 5. results of the m-svm polynomial test evaluation m-svm(polynomial) split ratio accuracy precision recall data training testing train test 10% 90% 71.2346% 53.6462% 56.8519% 1080 120 20% 80% 71.5278% 54.3445% 57.2917% 240 960 30% 70% 71.7460% 53.2519% 57.6190% 360 840 40% 60% 75.7407% 62.1921% 63.6111% 480 720 50% 50% 77.4444% 64.9680% 66.1667% 600 600 60% 40% 78.7500% 67.4113% 68.1250% 720 480 70% 30% 80.7407% 70.9712% 71.1111% 840 360 80% 20% 85.5556% 77.9906% 78.3333% 960 240 90% 10% 85% 77.5126% 77.5% 120 1080 in (table 5) explains that the results of the evaluation of test data using the m-svm polynomial method obtained the highest accuracy score, namely at a split ratio of 80:20 with an accuracy score = 85.5556%, precision = 78.9906%, and recall = 78.3333%. and the results of the calculation of the confusion matrix m-svm polynomial with the highest value can be seen in (figure 10). vol.6, no.2, july 2025 | 73 figure 10. cm m-svm polynomial 80:20 table 6. m-svm linear test evaluation results m-svm(linear) split ratio accuracy precision recall data training testing test lati 10% 90% 71% 52.5762% 56.6219% 1080 120 20% 80% 71.1078% 54.4445% 57% 960 240 30% 70% 71.1905% 55.6062% 56.7857% 840 360 40% 60% 73.0556% 58.3109% 59.5833% 720 480 50% 50% 72.8889% 57.9054% 59.3333% 600 600 60% 40% 77.7778% 66.9101% 66.6667% 480 720 70% 30% 80% 69.8135% 70% 360 840 80% 20% 84.4444% 76.5757% 76.6667% 240 960 90% 10% 83.8889% 76.2121% 75.8333% 120 1080 in (table 6) it is explained that the results of the evaluation of the test data using the m-svm linear method obtained the highest accuracy score, namely at a split ratio of 80:20 with an accuracy score = 84.4444%, precision = 76.5757%, and recall = 76.6667%. high accuracy, precision, and recall at a ratio of 80:20 occur because the model has enough data for training (80% of data for training). and the results of the calculation of the m-svm linear confusion matrix with the highest value can be seen in (figure 11). figure 11 cm m-svm linear 80:20 vol.6, no.2, july 2025 | 74 table 7 gaussian m-svm test evaluation results m-svm(gausian) split ratio accuracy precision recall data training testing test lati 10% 90% 61.7778% 45.8540% 42.6667% 1080 120 20% 80% 61.7778% 45.9603% 46.7% 960 240 30% 70% 63.4286% 47.992% 45.1429% 840 360 40% 60% 67.4074% 50.0876% 51.1111% 720 480 50% 50% 67.4667% 49.7449% 51.2% 600 600 60% 40% 67.7778% 49.9405% 51.6667% 480 720 70% 30% 73.2593% 55.5506% 56.8889% 360 840 80% 20% 76.8889% 58.8763% 59.3333% 240 960 90% 10% 79.8889% 64.7186% 65.3333% 120 1080 in (table 7) it is explained that the results of the evaluation of the test data using the gausian msvm method obtained the highest accuracy score, namely at a split ratio of 90:10 with an accuracy score = 69.7778%, precision = 52.5284%, and recall = 54.6667%. and the results of the calculation of the gausian m-svm confusion matrix with the highest value at a split ratio of 90:10 can be seen in (figure 12) figure 12. cm m-svm gaussian 90:10 d. results of comparison of training accuracy of multiclass support vector machine (m-svm) the accuracy results of the multiclass support vector machine (m-svm) method training are shown in the comparison in (table 8). the table provides an overview of how effective the m-svm method is in classifying data, and shows the variation in performance based on the composition of the data used. this allows for evaluating the advantages and disadvantages of the method in different contexts. vol.6, no.2, july 2025 | 75 table 8 results of training accuracy data comparison split ratio accuracy polynomial linear gaussian 10;90 88% 87.2222% 80% 20;80 89.7222% 89.4444% 79.2125% 30;70 85.1852% 84.8148% 79% 40;60 85.5556% 86.2500% 77.2222% 50;50 85.7778% 84.4444% 76% 60;40 85.2778% 84.9074% 74.6296% 70;30 85.2381% 84.2063% 71.1905% 80;20 84.5833% 84.5833% 71.5278% 90;10 85.1852% 84.3827% 71.2346% based on the results of the accuracy comparison in (table 8), it can be concluded that the method with the polynomial kernel shows the highest accuracy value of 89.7222% at a ratio of 20:80, and in general the performance of the polynomial kernel is superior to the linear and gaussian kernels. the linear kernel recorded the highest accuracy value of 88.4444% at a ratio of 20:80, while the gaussian kernel had the lowest performance, with the highest accuracy of only 80% at a ratio of 10:90. overall, polynomial kernels are more effective in handling larger training data, while linear kernels show better results at more balanced data ratios between training and testing data. gaussian kernels, although inferior, still provide good performance at more dominant testing data ratios, but not as high as polynomial and linear kernels. e. results of comparative accuracy of multiclass support vector machine (m-svm) testing the accuracy results of the testing data from the multiclass support vector machine (m-svm) method are shown in the comparison results in (table 9). the table provides an overview of how effective the m-svm method is in classifying data, and shows variations in performance based on the composition of the data used. table 9 results of comparison of test accuracy data split ratio accuracy polynomial linear gaussian 10;90 71.2346% 71% 61.7778% 20;80 71.5278% 71.1078% 61.7778% 30;70 71.7460% 71.1905% 63.4286% 40;60 75.7407% 73.0556% 67.4074% 50;50 77.4444% 72.8889% 67.4667% 60;40 78.7500% 77.7778% 67.7778% 70;30 80.7407% 80% 73.2593% 80;20 85.5556% 84.4444% 76.8889% 90;10 85% 83.8889% 79.8889% vol.6, no.2, july 2025 | 76 based on the results of the accuracy comparison in (table 9), it can be concluded that the methods with polynomial kernel and linear kernel show better performance compared to the gaussian kernel. the polynomial kernel has the highest accuracy value of 85.5556% at a ratio of 80:20, while the linear kernel achieves the highest accuracy value of 84.4444% at a ratio of 80:20. meanwhile, the gaussian kernel is recorded with the highest accuracy value of 79.8889% at a ratio of 90:10. overall, the polynomial kernel tends to be more effective in handling larger test data, with more stable accuracy across data ratios. the linear kernel, although slightly lower, still shows consistent results, while the gaussian kernel produces lower accuracy across almost all test data ratios. this suggests that the polynomial and linear kernels are more suitable for this rice leaf disease classification than the gaussian kernel iv. conclusions this study successfully built a rice leaf disease classification system using the multi-class support vector machine (m-svm) method. the test results showed that the polynomial kernel provided the highest accuracy of 85.56% at a training and testing ratio of 80:20, followed by the linear kernel (84.44%) and gaussian (79.89%). glcm and rgb-based feature extraction proved effective in supporting model performance through leaf texture and color analysis. references [1] s. sulistiyanto, ta saputri, and n. noviyanti, “early detection of rice pests and diseases using the certainty factor method,” jurikom (jurnal ris. komputer) , vol. 9, no. 1, p. 48, 2022, doi: 10.30865/jurikom.v9i1.3778. [2] ulfah nur oktaviana, ricky hendrawan, alfian dwi khoirul annas, and galih wasis wicaksono, “rice disease classification based on leaf images using resnet101 trained model,” j. resti (information systems and technology engineering) , vol. 5, no. 6, pp. 1216– 1222, 2021, doi: 10.29207/resti.v5i6.3607. [3] m. khoiruddin, a. junaidi, and wa saputra, “rice leaf disease classification using convolutional neural network,” j. dinda data sci. inf. technol. data anal. , vol. 2, no. 1, pp. 37–45, 2022, doi: 10.20895/dinda.v2i1.341. [4] r. naa et al. , “classification of papuan batik cloth motifs using the support vector machine (svm) 1,2 method,” vol. 11, no. 35, 2024. [5] s. andayani, “bacterial leaf blight disease,” bbpp lembang. accessed: nov. 02, 2024. [online]. available: https://bbpplemembang.bppsdmp.pertanian.go.id/publikasi-detail/1145 [6] e. nahak et al. , “disease classification in apple plants through leaf images using multiclass support vector machine 1,2 method,” vol. 11, no. 3, pp. 401–408, 2024. [7] f. tampinongkol, “identification of tomato leaf diseases using gray level co-occurrence matrix (glcm) and support vector machine (svm),” techno xplore j. comput. and technol. inf. , vol. 8, no. 1, pp. 08–16, 2023, doi: 10.36805/technoxplore.v8i1.3578. [8] nurul mudhofar and soffiana agustin, “classification of apple leaf diseases using rgb color feature extraction,” repeater publ. tech. inform. and jar. , vol. 2, no. 3, pp. 147–156, 2024, doi: 10.62951/repeater.v2i3.120. [9] c. wijaya, h. irsyad, and w. widhiarso, “pneumonia classification using k-nearest neighbor method with glcm extraction,” j. algoritm. , vol. 1, no. 1, pp. 33–44, 2020, doi: vol.6, no.2, july 2025 | 77 10.35957/algoritme.v1i1.431. [10] j.e. bata et al. , “on google maps using multi-class,” vol. 8, no. 6, pp. 11115– 11123, 2024. [11] kurniawan, i., hananto, a. l., hilabi, s. s., hananto, a., priyatna, b., & rahman, a. y. (2023). perbandingan algoritma naive bayes dan svm dalam sentimen analisis marketplace pada twitter. jatisi (jurnal teknik informatika dan sistem informasi), 10(1), 731-740. [12] priyatna, b., & hilabi, s. s. (2025). klasifikasi sentimen analisis ulasan aplikasi alfagift menggunakan algoritma long short term memory. storage: jurnal ilmiah teknik dan ilmu komputer, 4(2), 48-55. [13] priyatna, b., rahman, t. k. a., hananto, a. l., hananto, a., & rahman, a. y. (2024). mobilenet backbone based approach for quality classification of straw mushrooms (volvariella volvacea) using convolutional neural networks (cnn). joiv: international journal on informatics visualization, 8(3-2), 1749-1754. [14] salsabila, s. a., priyatna, b., & hananto, a. (2025). komparasi kinerja model naive bayes, svm, dan bert dalam klasifikasi sentimen ulasan pada aplikasi yummy. storage: jurnal ilmiah teknik dan ilmu komputer, 4(2), 42-47. [15] hananto, a. l., hananto, a., huda, b., rahman, a. y., novalia, e., & priyatna, b. (2024). determination of training participants in community work training centers using the naïve bayes classifier algorithm. joiv: international journal on informatics visualization, 8(3), 1162-1167. vol. 5, no.1 januari 2024 | 19 examining healthcare profesional’s acceptance of electronic medical records system using extended utaut2 milenia ayukharisma1, dian budi santoso2 1vocational school, gadjah mada university 2department of health services and information, vocational school, gadjah mada university 1mileniaayukharisma@gmail.com, 2dianbudisantoso@ugm.ac.id abstract this study aims to analyze user acceptance of the electronic medical record system using the extended unified theory of acceptance and use technology 2 model at pku muhammadiyah bantul hospital. the utaut 2 model was chosen because it is the latest technology acceptance model which is a unification, synthesis, or summary of the eight pre-existing technology acceptance models. the subjects in this study were pku muhammadiyah bantul hospital employees who used an electronic medical record system specifically for outpatient care. the object of this study is user acceptance of using the rme system in health services. data collection techniques in this study are using questionnaires and observation. this research is a type of quantitative analytic research with data analysis using descriptive analysis. data processing in this study used smart-pls software version 4.0 with sem-pls data analysis. the results showed that the aspects of the extended utaut2 model that had a positive and significant effect on user acceptance were performance expectancy (t=1.816), while the aspects of effort expectancy (t=0.419), social influence (t=0.635), facilitating conditions (t=0.139), hedonic motivation (t=0.909), price value (t=.304) habit (t=1.458), trust (t=0.032) and perceived risk (1.365) have no effect on user acceptance of the emr system. gender and age moderation variables were found to have no effect on the relationship between variables. keywords: emr, hospital, user, utaut2 i. introductions technology is developing rapidly in the current era of globalization. advances in technology have penetrated into various fields including the health sector. hospitals are required to build on quality of health services by utilizing currently developing technology. one of the implementations of technological advances in the health sector is the application of emr in health services [1]. emr is an electronic record that includes information such as a person's health that is created, collected, managed, used, and purchased by a doctor or health worker who is entitled to a health care institution. [1]. quality electronic medical records will produce optimal patient health services and produce complete information to support organizational or hospital decision making [2]. quality health services are supported by an optimal information system design [3]. in addition, user acceptance of the use of a system needs to be measured to produce a system that meets user needs [4]. user perceptions can help provide the right recommendations to maximize the development of electronic medical record systems [5]. one of the theory that can be used to measure the level of user acceptance of a system is using utaut2 theory [6]. utaut 2 was chosen to be used because this theory is the latest theory regarding user acceptance which adds new variables to support research in looking at acceptance of the use of new technology [7]. utaut2 was developed based on the utaut model which has four main constructs including 1) performance expectancy, 2) effort expectancy, 3) social influence, and 4) facilitating conditions which are then added again the three main constructs to support more accurate research include 1) hedonistic motivation, 2) price value, and 3) habit. through utaut2 it can be understood that users' reactions and perceptions of a technology can influence their attitude in accepting and using technology [23]. a technology can increase productivity optimally if users can accept the use of a technology and according to user needs [8]. in this study also adds new variables to utaut 2 to optimize research results that p-issn : 2715-2448 | e-issn : 2715-7199 vol.5 no.1 januari 2024 buana information technology and computer sciences (bit and cs) mailto:dian%20budi%20santoso@ugm.ac.id vol. 5, no.1 januari 2024 | 20 are appropriate in the field. the added variables are trust and perceived risk. there are moderator variables in this study, namely gender and age. pku muhammadiyah bantul hospital is a private hospital that has implemented electronic medical records since 2018. the emr system used was designed independently by the hospital. currently emr is implemented in outpatient services and continues to be developed in inpatient services. electronic medical record system at pku muhammadiyah bantul hospital still has various kinds of obstacles, including system errors, duplication of data entry, incomplete data output, patient data missing from the system without a known cause, and other obstacles that impede health services provided to patients. this needs to be analyzed to produce an optimal electronic medical record system and according to user needs. a good electronic medical record system will make patient treatment more optimal due to the continuity of medical history data owned by the patient. ii. methods in conducting this research, to measure user acceptance of using new technology, the extended utaut2 research concept was used because it has variables that match the research being carried out. the extended utaut2 conceptual framework is as follows. figure 1. research concept framework based on the figure 1, it can be seen that there are 9 independent variables and one dependent variable, as well as two moderator variables in this study, namely gender and age that using in this research. a type of quantitative analytical research is used in this study. this study used cross-sectional research. this research was conducted from may to july 2023. the population in this study were all pku muhammadiyah bantul hospital staff who used the outpatient emr system in health services. total population in this study was 152 officers who work as doctors, nurses, medical recorders, laboratories, radiographers, pharmacists, and health insurance center officers. sample is the object under study and is considered to represent the entire population [8]. sampling is the process of taking several elements from the population under study to be sampled, and understanding the various characteristics or characteristics of the subjects being sampled, which later can be generalized from the population elements [3]. proportional stratified random sampling was used in this study. this technique is used for populations that have heterogeneous and proportionally stratified members or elements [21]. utilizing the slovin formula to determine the number of samples as shown below. 𝑛 = 𝑁 1 + 𝑁 (𝑒)2 vol. 5, no.1 januari 2024 | 21 n = size of sample n = size of population e = percent allowance for sampling error that is acceptable for inaccuracy. calculation of the percent allowance in this study uses a percent of 10% and the sample calculation results are obtained as follows. 𝑛 = 152 1 + 152 (0,1)2 𝑛 = 152 1 + 152 (0,01) 𝑛 = 152 1 + 1,52 𝑛 = 152 2,52 𝑛 = 60 based on the above calculation, a minimum of 60 samples must be taken. the research instrument is a tool that is observed [21]. the research instruments used in this study were questionnaires and observation sheets. questionnaires are data collection techniques in the form of statements or questions given to respondents to answer [21]. the questionnaire used a likert scale of 1 to 5. the likert scale on the questionnaire is used to analyze the items in this variable research. questionnaire was distributed to employees of pku muhammadiyah bantul who utilized an electronic medical record system for data collection. in this study, the processes of data analysis were sem-pls analysis and univariate analysis. to determine the tendency of respondents' responses to the statement items on the research questionnaire, univariate was used to describe the characteristics of the respondents and display the distribution of each variable. pls is a variance based structural equation analysis that can test structural models and measurement models simultaneously. iii. results and discussions 1. the results of the description of the characteristics of the respondents characteristics respondents were divided based on gender, age, education background, and the profession of the respondents. respondents in this study amounted to 60 respondents with the details of the respondents as shown in the following table. table 1. gender of respondents gender totals female 42 male 18 the ages of the respondents in this study were grouped into five age groups, as follows. table 2. ages range of respondents ages range totals 17-25 years old 12 26-35 years old 18 36-45 years old 10 46-55 years old 14 >56 years old 6 the results of the study show that the last educational background of the respondents is as follows. table 3. educational background of respondents graduate totals senior high school 1 diploma iii 31 vol. 5, no.1 januari 2024 | 22 bachelor 10 masters 18 respondents involved in this study shown in the following table. table 4. profession of respondent profession totals doctors 18 nurses 9 medical recorders 12 radiographers 5 pharmacists 4 laboratories 7 health insurance officers 5 2. sem-pls analysis results the outer model test and the inner model test are the two methods that used in partial least square testing. a. outer model (measurement model) 1) convergent validity test table 5. result of convergent validity variabel item loading factor ave performance expectancy pe1 0,804 0,536 pe2 0,680 pe3 0,570 pe4 0,841 effort expectancy ee1 0,749 0,630 ee2 0,914 ee3 0,882 ee4 0,590 social influence si1 0,914 0,846 si2 0,957 si3 0,887 facilitating condition fc1 0,739 0,601 fc2 0,716 fc3 0,763 fc4 0,873 hedonic motivation hm1 0,882 0,726 hm2 0,902 hm3 0,766 price value pv1 0,726 0,627 pv2 0,714 pv3 0,919 habit h1 0,806 0,604 h2 0,717 h3 0,836 h4 0,745 trust t1 0,823 0,623 t2 0,768 t3 0,754 t4 0,810 perceived risk pr1 0,884 0,722 pr2 0,844 pr3 0,821 acceptance and use of technology aus1 0,949 0,894 aus2 0,958 aus3 0,930 gender zg 1,000 1,000 vol. 5, no.1 januari 2024 | 23 variabel item loading factor ave age zu 1,000 1,000 table 5 shows the loading factor value of each item on the variable has a value of > 0,5 so that the instrument is declared valid according to convergent validity testing as well as the ave value of each construct in the entire model has a value of > 0,5 so it can be concluded that all constructs are valid and fulfill validity converge well. 2) discriminant validity test figure 2. discriminant validity test result the results in this study in figure 2 show that all constructs have a construct correlation with measurement items that is higher than the other constructs. it means that the requirements for a good discriminant validity test have been fulfilled. aus ee fc h hm pe pr pv si t zg zu aus1 0.949 0.628 0.572 0.513 0.556 0.579 -0.351 0.309 0.523 0.603 0.029 -0.125 aus2 0.958 0.581 0.599 0.568 0.651 0.541 -0.336 0.322 0.576 0.630 0.073 -0.102 aus3 0.930 0.511 0.485 0.613 0.639 0.461 -0.274 0.416 0.496 0.582 -0.034 -0.072 ee1 0.362 0.749 0.370 0.445 0.211 0.279 -0.267 0.017 0.321 0.443 -0.041 -0.079 ee2 0.477 0.914 0.377 0.305 0.392 0.424 -0.262 0.136 0.281 0.467 -0.008 -0.073 ee3 0.635 0.882 0.538 0.405 0.381 0.352 -0.262 0.143 0.241 0.534 0.083 -0.251 ee4 0.376 0.590 0.444 0.032 0.440 0.319 -0.191 0.376 0.368 0.324 0.067 -0.118 fc1 0.330 0.365 0.739 0.359 0.249 0.282 -0.110 0.227 0.209 0.455 0.260 -0.088 fc2 0.360 0.442 0.716 0.365 0.296 0.273 -0.201 0.219 0.322 0.482 0.118 -0.113 fc3 0.521 0.478 0.763 0.151 0.664 0.393 -0.205 0.494 0.294 0.392 0.115 -0.118 fc4 0.540 0.422 0.873 0.343 0.568 0.426 -0.192 0.454 0.367 0.594 0.157 -0.099 h1 0.626 0.458 0.449 0.806 0.282 0.222 -0.196 0.074 0.341 0.568 -0.072 -0.190 h2 0.382 0.247 0.034 0.717 0.270 0.077 0.076 0.220 0.415 0.456 -0.006 0.029 h3 0.404 0.243 0.307 0.836 0.221 0.141 -0.077 0.115 0.419 0.541 -0.000 -0.017 h4 0.342 0.149 0.278 0.745 0.103 0.092 -0.034 0.244 0.339 0.533 -0.020 -0.038 hm1 0.546 0.420 0.551 0.156 0.882 0.507 -0.208 0.505 0.405 0.402 0.166 0.023 hm2 0.666 0.400 0.512 0.444 0.902 0.594 -0.271 0.472 0.535 0.564 0.077 0.088 hm3 0.407 0.326 0.524 0.075 0.766 0.373 -0.106 0.592 0.413 0.329 0.159 0.224 pe1 0.431 0.397 0.367 0.234 0.308 0.804 -0.023 0.094 0.279 0.324 0.028 0.129 pe2 0.386 0.045 0.276 0.240 0.451 0.680 -0.008 0.380 0.383 0.325 -0.061 0.234 pe3 0.326 0.342 0.189 -0.106 0.413 0.570 -0.073 0.311 0.248 0.110 0.023 0.276 pe4 0.475 0.461 0.461 0.138 0.558 0.841 -0.103 0.310 0.303 0.276 -0.006 0.134 pr1 -0.322 -0.270 -0.126 -0.038 -0.272 -0.152 0.884 -0.254 -0.226 -0.381 -0.052 0.258 pr2 -0.215 -0.262 -0.229 -0.132 -0.056 -0.011 0.844 -0.088 -0.060 -0.329 0.004 0.423 pr3 -0.305 -0.255 -0.253 -0.101 -0.241 0.001 0.821 -0.008 -0.050 -0.287 0.007 0.528 pv1 0.158 0.263 0.443 0.242 0.278 0.140 -0.116 0.726 0.257 0.473 0.224 0.057 pv2 0.166 0.092 0.245 0.222 0.255 0.290 -0.146 0.714 0.531 0.391 0.217 0.160 pv3 0.418 0.165 0.443 0.101 0.674 0.370 -0.111 0.919 0.366 0.383 0.103 0.054 si1 0.505 0.383 0.414 0.429 0.555 0.418 -0.151 0.466 0.914 0.519 0.105 0.103 si2 0.532 0.378 0.323 0.455 0.460 0.343 -0.175 0.398 0.957 0.456 0.100 -0.001 si3 0.516 0.242 0.345 0.438 0.463 0.379 -0.062 0.398 0.887 0.467 0.141 -0.016 t1 0.624 0.486 0.546 0.675 0.404 0.335 -0.370 0.279 0.506 0.823 0.107 -0.234 t2 0.323 0.314 0.383 0.646 0.268 0.123 -0.282 0.329 0.377 0.768 0.148 -0.150 t3 0.522 0.505 0.509 0.332 0.621 0.442 -0.233 0.568 0.371 0.754 0.085 -0.074 t4 0.460 0.429 0.469 0.508 0.296 0.159 -0.339 0.358 0.361 0.810 0.099 -0.239 zg 0.025 0.040 0.199 -0.040 0.148 -0.006 -0.019 0.186 0.125 0.134 1.000 -0.023 zu -0.105 -0.179 -0.135 -0.093 0.117 0.250 0.468 0.094 0.030 -0.225 -0.023 1.000 vol. 5, no.1 januari 2024 | 24 3) reliability test table 6. reliability test variabel cronbach’s alpha composite reliability (rho_a) composite reliability (rho_c) result acceptance and use of technology 0,941 0,942 0,962 reliable effort expectancy 0,795 0,857 0,869 reliable facilitating condition 0,783 0,810 0,857 reliable habit 0,789 0,832 0,859 reliable hedonic motivation 0,814 0,860 0,888 reliable performance expectancy 0,701 0,728 0,819 reliable perceived risk 0,810 0,830 0,886 reliable price value 0,744 1,056 0,833 reliable social influence 0,908 0,909 0,943 reliable trust 0,802 0,824 0,869 reliable zg 1,000 1,000 1,000 reliable zu 1,000 1,000 1,000 reliable based on table 6, all variables have met the reliability requirements, that is all variables have a value of > 0,7 so that the measurement model used in this study can be declared reliable. b. inner model (structural model) 1) coefficient of determination table 7. coefficient determination result variable (r2) chategory aus 0,756 strong the results based on table 7 show that (r2) value on the dependent variable aus of 0,712, meaning that this explains that all variables have an influence of 75,6% on acceptance and use of technology. 2) hypothesis test table 8. hypothesis test result t statistic p values result h1 pe→ aus 1.816 0.035 accepted h2 ee → aus 0.419 0.338 rejected h3 si → aus 0.635 0.263 rejected h4 fc → aus 0.139 0.445 rejected h4a fc*zg→aus 0.290 0.386 rejected h4b fc*zu→aus 0.691 0.245 rejected h5 hm → aus 0.909 0.182 rejected h5a hm*zg→aus 0.292 0.385 rejected h5b hm*zu→aus 0.389 0.349 rejected h6 pv → aus 0.304 0.381 rejected h6a pv*zg→aus 0.713 0.238 rejected h6b pv*zu→aus 0.231 0.409 rejected vol. 5, no.1 januari 2024 | 25 t statistic p values result h7 h → aus 1.458 0.073 rejected h7a h*zg→aus 0.153 0.439 rejected h7b h*zu→aus 0.407 0.342 rejected h8 t → aus 0.032 0.487 rejected h9 pr → aus 1.365 0.086 rejected h1: performance expectancy have a positive and significant effect on user acceptance of the electronic medical record system the results of data processing obtained a t value of 1.816 (> 1.64) with a p value of 0.035 (< 0.05). based on the test results, hypothesis 1 is accepted. performance expectations have a positive and significant effect on user acceptance of the rme system. the results of this study are in line with [14], and [15]. which state that performance expectations affect user acceptance. the results of this study were also reinforced by researchers [12]. that this research shows that performance expectancy affects the acceptance and use of systems or technology. h2: effort expectancy have a positive and significant effect on user acceptance of the electronic medical record system the results of data processing obtained a t value of 0.419 (<1.64) and a p value of 0.338 (>0.05). hypothesis 2 is rejected, effort expectations have no effect on user acceptance of the emr system. the results of this study are in line with research [11] which states that effort expectations have no effect on user acceptance of the system. the results of this study are also reinforced by research conducted [25] which states that effort expectations do not have a significant effect on user acceptance of systems or technology. h3: social influence has a positive and significant effect on user acceptance of the electronic medical record system the results of data processing obtained a t value of 0.635 (<1.64) and a p value of 0.263 (>0.05). the t value and the p value show that hypothesis 3 is rejected so that social influence does not affect user acceptance of the rme system. this research is in contrast to research [19] in which social influence has a significant influence on the behavioral intention to use the system. however, the results of this study are in line with previous research [5] which states that social influence has no influence on someone in using a system. h4: facilitating conditions have a positive and significant effect on user acceptance of the electronic medical record system the results of data processing obtained a t value of 0.139 (<1.64) and a p value of 0.445 (>0.05). based on the test findings, hypothesis 4 is rejected so that facilitating conditions do not affect user acceptance of the rme system. the results of this study are in line with research [2] which also found that facilitating conditions did not have a significant effect on behavioral interest in accepting the use of the system. h4a: gender moderates the effect of facilitating conditions on user acceptance of the electronic medical record system the results of data processing obtained t = 0.290 (<1.64) and p = 0.386 (> 0.05). the hypothesis is rejected so that gender has no effect in moderating conditions that facilitate user acceptance of the rme system. h4b: age moderates the effect of facilitating conditions on user acceptance of the electronic medical record system the results of data processing obtained t = 0.691 (<1.64) and p = 0.245 (> 0.05). the hypothesis is rejected so that age has no effect in moderating conditions that facilitate user acceptance of the rme system. h5: hedonistic motivation has a positive and significant effect on user acceptance of the electronic medical record system the results of data processing obtained a t value of 0.909 (<1.64) with a p value of 0.182 (>0.05). hypothesis 5 is rejected. hedonistic motivation has no effect on user acceptance vol. 5, no.1 januari 2024 | 26 of the rme system. this research is in line with research conducted (ismarmiaty and etmy, 2018) which states that hedonistic motivation has no effect on system user acceptance. the results of the study which stated that hedonism motivation had no influence on user acceptance were also strengthened by research conducted by [3] who had the same research results. h5a : gender moderates the effect of hedonistic motivation on user acceptance of the electronic medical record system the results of data processing obtained t = 0.292 (<1.64) and p = 0.385 (> 0.05). the hypothesis is rejected so that gender has no effect in moderating hedonistic motivation on user acceptance of the rme system. h5b: age moderates the influence of hedonistic motivation on user acceptance of the electronic medical record system the results of data processing obtained t = 0.389 (<1.64) and p = 0.349 (> 0.05). the hypothesis is rejected so that age has no effect in moderating hedonistic motivation on user acceptance of the rme system. h6: price value has a positive and significant effect on user acceptance of electronic medical records the results of data processing obtained a t value of 0.304 (<1.64) and a p value of 0.381 (> 0.05). based on the test findings, hypothesis 6 is rejected. the price value has no effect on user acceptance of the rme system. research of (ismarmiaty and etmy, 2018) has the same results, price value has no effect on system user acceptance. h6a: gender moderates the effect of price value on acceptance of electronic medical record users the results of data processing obtained t = 0.713 (<1.64) and p = 0.238 (> 0.05). the hypothesis is rejected so that gender has no effect in moderating the value of prices on user acceptance of the rme system. h6b: age moderates the effect of price value on user acceptance of the electronic medical record system the results of data processing obtained t = 0.231 (<1.64) and p = 0.409 (> 0.05). the hypothesis is rejected so that age has no effect in moderating the price value on user acceptance of the rme system. h7: habits have a positive and significant effect on user acceptance of the electronic medical record system the results of data processing obtained a t value of 1.458 (<1.64) and a p value of 0.073 (>0.05). hypothesis 7 is rejected, habit has no effect on user acceptance of the rme system. this results are in line with research [4] which suggests that habits have no effect on user acceptance of the system. the results of the study that habit has no effect on system user acceptance is also reinforced by research results [6] habit has no effect on system user acceptance. h7a: gender moderates the influence of habits on user acceptance of the electronic medical record system the results of data processing obtained t = 0.153 (<1.64) and p = 0.439 (> 0.05). the hypothesis is rejected so that gender has no effect in moderating habits on user acceptance of the rme system. h7b: age moderates the influence of habits on acceptance of users of the electronic medical record system the results of data processing obtained t = 0.407 (<1.64) and p = 0.342 (> 0.05). the hypothesis is rejected so that age has no effect in moderating habits on user acceptance of the rme system. h8: trust has a positive and significant effect on user acceptance of the electronic medical record system the results of data processing obtained a t value of 0.032 (<1.64) a p value of 0.487 (> 0.05). hypothesis 8 is rejected, trust has no effect on user acceptance of the rme system. the results of this study are in line with research [11] which also obtained the result that trust does not have a significant effect on user acceptance of the system. vol. 5, no.1 januari 2024 | 27 h9: perceived risk has a positive and significant effect on user acceptance of the electronic medical record system the results of data processing obtained a t value of 1.365 (<1.64) with a p value of 0.086 (<0.05). hypothesis 9 is rejected, meaning that risk perception has no effect on user acceptance of the rme system. this research is supported by research (anggraeni et al, 2023) which explains that perceptions of risk have no effect on user acceptance of systems or technology. the results of this study are reinforced by research conducted [22] and [8] which state that risk perception has no effect on system user acceptance. test analysis on the influence of moderating variables gender and age moderation has no effect in this study. this was also found in other studies where there were several factors that could cause the moderating variable not to affect the independent variable on the dependent variable. in research conducted by [10] it was also found that gender had no effect in moderating the independent variable on the dependent variable utaut2. [13] stated that gender and age have no effect due to the evolution of individuals into modern society. this makes there is no difference or no difference between gender and age level in using technology from the user's perspective. in terms of gender [7] states that gender roles can no longer be used as a benchmark in predicting the use of new technology. in terms of age, the reason why age does not moderate the relationship between variables is because there is a balanced age grouping of respondents where a balanced age distribution will better describe the gender effect. iv. conclusions in this study, based on the tests that have been carried out, interesting results are obtained that the factors that can influence a person to accept and use a system or technology are aspects of business expectations. based on this research, it can be concluded that in making a new innovation in a system or technology, developers must pay attention to aspects of user performance efficiency in order to increase work productivity. references [1] alazzam, m. b., basari, a. s. h., sibghatullah, a. s., doheir, m., enaizan, o. m. a., & mamra, a. h. k. (2015). ehrs acceptance in jordan hospitals by utaut2 model: preliminary result. journal of theoretical and applied information technology, 78(3), 473–482. [2] auliya, n. (2018). penerapan model unified theory of acceptance and, 1–10 [3] handayani, r. (2020). metodologi penelitian sosial. yogyakarta: trussmedia grafika. [4] hasibuan, h.t. (2021). faktor-faktor yang mempengaruhi minat menggunakan layanan financial technology peer to peer lending syariah. e-jurnal akuntansi, 31(5), 1201. [5] ismarmiaty, i. and etmy, d. (2018). model pendekatan utaut2 modifikasi pada analisis penerimaan dan penggunaan teknologi e-government di nusa tenggara barat. matrik :jurnal manajemen, teknik informatika dan rekayasa komputer, 18(1), 106–114. [6] mayanti, r. (2022). preferensi masyarakat terhadap quick response code indonesian standard sebagai sarana teknologi pembayaran digital. faktor exacta, 15(1), 1979–276. [7] moryson, h. and moeser, g. (2016).consumer adoption of cloud computing services in germany: investigation of moderating effects by applying an utaut model. international journal of marketing studies, 8(1), 14. [8] mubiyantoro, a. and syaefullah (2015). pengaruh persepsi kegunaan, persepsi kemudahan penggunaan, persepsi kesesuaian, dan persepsi risiko terhadap sikap penggunaan mobile banking. dk, 53(9), 1689–1699. [9] nastiti, i., & santoso, d. b. (2022). evaluasi penerapan sistem informasi manajemen rumah sakit di rsud slg kediri dengan menggunakan metode hot-fit. jurnal kesehatan vokasional, 7(2), 85. [10] nitcheu tcheuffa, p.c., kala kamdjoug, j.r. and fosso wamba, s. (2020) . moderating effects of age and gender on social commerce adoption factors the cameroonian context. lecture notes in information systems and organisation, 35(march), 263–274. [11] nuraini, d. (2020). analisis penerimaan aplikasi mobile tix id menggunakan model utaut vol. 5, no.1 januari 2024 | 28 2 extend anny mardjo. [skripsi]. uin syarif hidayatullah jakarta. [12] nurissobah, m. (2020). analisis user experience sapto oleh perguruan tinggi di indonesia. [thesis]. universitas gadjah mada. [13] palau-saumell, r. et al. (2019). user acceptance of mobile apps for restaurants: an expanded and extended utaut-2. sustainability, 11(4), 1210. [14] pertiwi, n.w.d.m.y. and ariyanto, d. (2021). penerapan model utaut 2 untuk menjelaskan minat dan perilaku penggunaan mobile banking. e-jurnal akuntansi, 31(10), 2569. [15] pramudiana, a. m. p. dan y. (2015). pengaruh faktor-faktor dalam modifikasi unified theory of acceptance and use of technology 2 terhadap perilaku konsumen dalam mengadopsi layanan wifi pt. xyz area jakarta, 2(2), 1085–1094. [16] pratama, m. h., & darnoto, s. (2017). analisis strategi pengembangan rekam medis elektronik di instalasi rawat jalan rsud kota yogyakarta. jurnal manajemen informasi kesehatan indonesia, 5(1), 34. [17] pujihastuti, a. (2021). penerapan sistem informasi manajemen dalam mendukung pengambilan keputusan manajemen rumah sakit. jurnal manajemen informasi kesehatan indonesia, 9(2), 200. [18] putra, g., & ariyanti, m. (2017). pengaruh faktor-faktor dalam modified unified theory of acceptance and use of technology 2 (utaut 2) terhadap niat prospective users untuk mengadopsi home digital services pt. telkom di surabaya. jurnal manajemen indonesia, 14(1), 59. [19] putra, m.a. (2018). evaluasi penggunaan pada produk uang elektronik e-money bank mandiri menggunakan modeul utaut 2 (studi kasus: kecamatan ciputat). energies, 6(1), pp. 1–8. [20] sari, a. p. (2019). pengukuran keberhasilan penerapan sistem institutional repository di uin syarif hodayatullah jakarta menggunakan human organization technology (hot) fit mode, 8(5), 55. [21] sugiyono. (2017) metode penelitian kuantitatif, kualitatif, dan r&d. bandung: alfabeta. [22] tazkiyatunnisa anggraeni, n., kresnamurti rivai p, a. and aditya, s. (2023). pengaruh perceived risk dan online customer review terhadap keputusan pembelian melalui kepercayaan pada pengguna marketplace di kota bekasi. sinomika journal: publikasi ilmiah bidang ekonomi dan akuntansi, 1(5), 1311–1322. [23] venkatesh, viswanath, james y. l. thong, xin xu. (2012). consumer acceptance and use of information technology: extendingthe unified theory of acceptance and use of technology. mis quarterly,36(1), 157-178. [24] wardani, r., tarbiati, u., fauziah, t. r., mahadewi, g. a. a. m., nahdlah, m. p., sudewa, i. g. n. w., & sakti, e. m. (2022). strategi pengembangan rekam medis elektronik di instalasi rawat jalan rsud gambiran kota kediri. madaniya pustaka, 3(1), 37–46. [25] widnyana, i.i.d.g.p. and yadnyana, i.k. (2015). implikasi model utaut dalam menjelaskan faktor niat dan penggunaan sipkd kabupaten tabanan. jurnal akuntansi universitas udayana, 112, 2302–8556. [26] widodo, m., irawan, m.i. and sukmono, r.a. (2019). extending utaut2 to explore digital wallet adoption in indonesia. 2019 international conference on information and communications technology, icoiact 2019, 878–883. vol. 6, no.2, july 2025 | 78 p-issn: 2715-2448 | e-issn: 2715-7199 vol.6 no.2 july 2025 buana information technology and computer sciences (bit and cs) development of a web-based youth innovation village system using laravel framework ratna suminar 1, lila setiyani 2*, ahmad mubarok 3 1, 2, 3 horizon university indonesia e-mail: ratna.suminar.stmik@krw.horizon.ac.id, lila.setiyani.krw@horizon.ac.id*, ahmad.mubarok.stmik@krw.horizon.ac.id received: 2025/03/04 | revised: 2025/07/14 | accepted: 2025/07/27 abstract youth play a crucial role in social development and serve as agents of change contributing to national progress. karang taruna, as a youth organization, has a strategic role in enhancing social welfare, including in mekarjati village. since its establishment on november 10, 2013, karang taruna mekarjati has grown with 13 subunits in each neighborhood unit (rw), serving a community of 11,881 people. however, the organization's business processes, such as member registration, social activity management, leadership training, and the formation and monitoring of msmes or startups, are still conducted manually. this results in operational inefficiencies and increases the risk of administrative errors. this study aims to design and develop a web-based system called "youth innovation village" to support the digitalization of karang taruna mekarjati’s management. the system is built using the laravel framework, with process modeling based on business process modeling notation (bpmn), the php programming language, and the mysql database. the software development approach used is agile scrum, allowing flexibility in adapting to the organization's needs throughout the system development life cycle. the results of this study indicate that the implementation of youth innovation village significantly enhances the effectiveness and efficiency of karang taruna’s management. the system accelerates membership administration, optimizes social activity coordination, and provides direct benefits to the organization's members in mekarjati village. through this digitalization, karang taruna can become more adaptive to technological advancements and expand its social impact more effectively. keywords: youth innovation village, laravel, bpmn, agile scrum, digitalization, operational efficiency i. introduction youth represent the power and strength of a nation, signifying that the younger generation plays a crucial role in national development. as social revolutionaries, youth are considered strategic agents of change due to their strong mentality, high potential, strong competitiveness, and quick-thinking ability [1] . given this vital role, karang taruna was established as a social organization to raise awareness and responsibility, particularly among young people, from the regional to the urban level. the primary focus of this organization is on social welfare rather than profit-making [5]. the mekarjati village government, located in west karawang district, karawang regency, has established karang taruna mekarjati, which has a significant role in fostering youth empowerment and community development. since its founding on november 10, 2013, karang taruna mekarjati has grown into a structured organization consisting of 13 subunits in each rw (rukun warga), serving a total population of 11,881 people. the organization carries out several key business processes, including member registration, social activity implementation, leadership training programs, and the formation and monitoring of msmes (micro, small, and medium enterprises) and startups [3]. despite its significant contributions, karang https://www.zotero.org/google-docs/?foyo1a https://www.zotero.org/google-docs/?f8a4ww vol. 6, no.2, july 2025 | 79 taruna mekarjati faces challenges in managing its operations due to the absence of an integrated digital system. most of its administrative processes, such as membership registration, proposal and execution of social activities, leadership training records, and business development monitoring, are still conducted manually. this manual approach results in inefficiencies, increased risk of errors, and difficulties in tracking the progress of various initiatives [2]. digital transformation in non-profit youth organizations has been recognized as a key enabler for improving operational effectiveness, transparency, and community engagement (digital transformation, n.d.). therefore, the need for a digital information system that can streamline and optimize karang taruna’s business processes has become increasingly urgent. to address this issue, this study aims to design and develop a web-based system called "youth innovation village" to be implemented in karang taruna mekarjati. the system will support the organization's business processes, ensuring improved efficiency and effectiveness in its operations. the development of this system adopts the uml (unified modeling language) modeling approach, bpmn (business process modeling and notation), php (hypertext preprocessor) as the programming language, mysql as the database management system, and the laravel framework [6]. this research is expected to contribute significantly to karang taruna mekarjati by enhancing the efficiency and effectiveness of its business processes. additionally, it will provide direct benefits to karang taruna members by enabling better management of activities, improving data accuracy, and facilitating easier access to organizational resources. with the implementation of the youth innovation village system, karang taruna mekarjati can leverage digital technology to maximize its impact on the local community, ensuring sustainable youth development and improved organizational governance. youth play a strategic role in social development as agents of change who drive innovation and societal transformation. according to chigunta (2002), young people have great potential in shaping innovation ecosystems through various community-based initiatives [1]. furthermore, the united nations development programme (2020) report emphasizes that youth participation in social activities can improve community well-being and build their leadership capacities for the future [11]. in indonesia, the karang taruna organization plays an essential role in increasing youth participation in social development. as a social youth organization, karang taruna focuses on social welfare development and youth empowerment at the village and sub-district levels [4] . digital transformation has become a key factor in enhancing the effectiveness of social organizations. drucker (2019) highlights that organizations that fail to adapt to technology tend to experience operational inefficiencies [2]. a study by goyal and sergi (2021) indicates that adopting digital technology in non-profit organizations can improve transparency, administrative efficiency, and member participation in organizational activities [8],[15]. however, many social organizations in indonesia still struggle to implement an integrated information system, leading to manual business processes that hinder efficiency and accuracy. laravel is one of the most popular php frameworks used in web-based system development. sommerville (2020) states that laravel offers modern features such as model-view-controller (mvc) architecture, object-relational mapping (orm), and enhanced security, which facilitate developers in building robust and structured applications [9]. fowler (2021) further explains that the use of unified modeling language (uml) and business process model and notation (bpmn) in software development helps in understanding business processes and system requirements effectively [6], [13]. agile scrum is a widely adopted software development methodology that emphasizes flexibility and quick iterations to meet user needs. schwaber & sutherland (2020) state that this method is highly effective for developing systems that require continuous adaptation and real-time improvements[10], [14]. several studies suggest that agile scrum enhances collaboration among development teams and ensures that the software produced aligns with organizational needs [3]. therefore, this method is https://www.zotero.org/google-docs/?8c0exq https://www.zotero.org/google-docs/?4u6lly https://www.zotero.org/google-docs/?4u6lly https://www.zotero.org/google-docs/?4u6lly https://www.zotero.org/google-docs/?e0kwav https://www.zotero.org/google-docs/?fd2xta https://www.zotero.org/google-docs/?d7aci5 https://www.zotero.org/google-docs/?d5dxo9 vol. 6, no.2, july 2025 | 80 chosen in this research to develop the youth innovation village system to be more responsive to the needs of karang taruna mekarjati. based on the literature review, it is evident that youth involvement in social development, challenges in digitization for social organizations, the use of laravel as a software development framework, and the implementation of agile scrum as a development methodology form a strong foundation for the development of the youth innovation village system. by referring to previous studies, this research is expected to contribute significantly to the digitalization of social organizations, particularly in the management of karang taruna mekarjati [12]. ii. methods this research employs a qualitative and quantitative approach to develop a web-based system, "youth innovation village," for the karang taruna mekarjati organization. the study follows the software development life cycle (sdlc) using the agile scrum methodology, incorporating system modeling with unified modeling language (uml) and business process model and notation (bpmn) to ensure an efficient and adaptable system. this study is applied research aimed at developing a digital solution to optimize the business processes of karang taruna mekarjati. the research stages consist of: figure 1. research stages 1. problem identification: a. analysis of current business processes at karang taruna mekarjati. b. identification of inefficiencies in membership registration, event management, leadership training, and msme/startup monitoring. this is an example of a youth organization (karang taruna) existing business process using bpmn in indonesian. figure 2. existing business process for karang taruna member registration problem identification system requirement analysis system design system development testing and evaluation implementation and deployment vol. 6, no.2, july 2025 | 81 figure 3. existing business processes for implementing karang taruna activities figure 4. existing business process of cadre development and leadership training figure 5. existing business process of umkm formation or startup monitoring 2. system requirement analysis: a. requirement gathering through interviews and observation with karang taruna members and administrators. b. documentation of functional and non-functional requirements. 3. system design: a. system architecture design using laravel framework. b. process modeling using bpmn for business workflow representation. c. system structure modeling using uml diagrams (use case, activity, and class diagrams). this is an example of uml diagram in this application. vol. 6, no.2, july 2025 | 82 figure 6. use case diagram vol. 6, no.2, july 2025 | 83 figure 7. activity diagram vol. 6, no.2, july 2025 | 84 figure 8. sequence diagram figure 9. class diagram vol. 6, no.2, july 2025 | 85 4. system development: a. implementation of the laravel framework for back-end development. b. php (hypertext preprocessor) for dynamic web processing. c. mysql as the relational database management system. d. html, css, and javascript for front-end interface design. 5. testing and evaluation: a. unit testing: testing individual components of the system. b. integration testing: verifying interactions between different modules. c. user acceptance testing (uat): evaluating the system with karang taruna members and administrators. d. performance testing: assessing system response time and efficiency. figure 10. ui beranda(home) 6. implementation and deployment: a. hosting the system on a cloud server for accessibility. b. conducting user training for karang taruna administrators and members. c. continuous monitoring and maintenance for long-term usability. iii. results and discussions this section presents the findings from the development and implementation of the "youth innovation village" system. the results are evaluated based on system functionality, performance, and user feedback. 1. system implementation and features the "youth innovation village" system was successfully developed and implemented using the laravel framework, php, mysql, and bpmn-based process modeling. the system was deployed on a cloud server to ensure accessibility and scalability. the main features of the system include: a. membership registration module 1) users can register online and receive digital membership ids. 2) admins can verify and manage member data efficiently. vol. 6, no.2, july 2025 | 86 figure 6. screenshot of membership registration module b. event management system 1) organizes and tracks social events, leadership training, and business development programs. 2) members can sign up for activities through the platform. figure 7. screenshot of event management module c. startup and msme monitoring module 1) provides a dashboard for tracking startup progress and msme activities. 2) enables mentorship and networking opportunities. vol. 6, no.2, july 2025 | 87 figure 8. screenshot of start-up and msme monitoring module d. administrative dashboard 1) allows karang taruna administrators to oversee all organizational activities. 2) provides real-time analytics and report generation for decision-making. e. communication and announcement system 1) facilitates direct communication between members and administrators. 2) features an announcement board for important updates. 2. system performance evaluation a. functional testing the system underwent unit testing, integration testing, and user acceptance testing (uat). results showed: table 1. functional testing test type success rate (%) issues encountered unit testing 98% minor ui alignment issues integration testing 95% some delays in data synchronization user acceptance testing (uat) 93% users requested additional search filters the membership registration and event management modules functioned smoothly, demonstrating seamless user interactions and efficient data processing. users found the registration process intuitive, allowing for quick and hassle-free onboarding. additionally, the startup monitoring feature was well-received, as it provided valuable insights into business development within karang taruna mekarjati. however, feedback indicated a need for additional filtering options to enhance data retrieval and streamline navigation. furthermore, some ui improvements were suggested to enhance accessibility and user experience, ensuring that all members, including those less familiar with digital platforms, could efficiently utilize the system. b. performance testing the system’s performance was measured based on response time, scalability, and concurrent user handling. tabel 2. performance testing parameter expected value measured value page load time < 2 seconds 1.5 seconds concurrent users 100+ users 150 users successfully handled vol. 6, no.2, july 2025 | 88 data processing time < 3 seconds 2.2 seconds the system performed within acceptable parameters, ensuring fast response times and smooth user interactions. 3. user feedback and satisfaction a survey was conducted among karang taruna mekarjati members and administrators after system deployment. 85 respondents participated, providing the following feedback: tabel 3. user feedback and satisfaction evaluation criteria satisfaction level (%) ease of use 92% system responsiveness 89% feature relevance 87% overall satisfaction 91% users highly appreciated the intuitive interface and the ease of navigation, which allowed them to efficiently interact with the system without requiring extensive technical knowledge. the wellstructured layout and user-friendly design contributed to a positive experience, enabling seamless access to various features. additionally, administrators found the dashboard particularly useful for decision-making, as it provided real-time data and analytics, helping them monitor membership activities, events, and startup developments more effectively. however, some respondents suggested the addition of a mobile app version to improve accessibility, ensuring that users could engage with the platform conveniently from their smartphones, particularly those who rely more on mobile devices than desktop computers. 4. discussion: comparison with manual system before digitalization, karang taruna mekarjati relied on manual record-keeping, paper-based registrations, and fragmented communication, leading to inefficiencies. the new system streamlined operations and introduced real-time monitoring, data automation, and centralized information management. table 4. comparison with manual system manual system digital system (youth innovation village) membership management paper-based, slow processing online, instant verification event registration manual sign-ups, prone to errors automated, real-time updates msme & startup tracking limited visibility digital dashboards & analytics communication whatsapp/manual calls centralized announcement system data storage physical records, high risk of loss secure cloud-based database the findings align with previous research on digital transformation in social organizations (goyal & sergi, 2021; lee & kotler, 2022), reinforcing that technology adoption improves efficiency, transparency, and engagement. 5. challenges and limitations despite its success, the system faced some challenges: vol. 6, no.2, july 2025 | 89 a. user adaptation issues: some members, especially older users, required training to navigate the platform. b. limited initial features: additional functionalities, such as mobile support and financial tracking, were suggested by users. c. internet dependency: the system requires a stable internet connection, which could be a limitation in remote areas. 6. future work to address current limitations and further improve the system, future research and development should focus on: a. mobile application development: to increase accessibility and engagement. b. financial management features: for tracking donations and event budgets. c. machine learning for data insights: to predict event attendance and optimize resource allocation. iv. conclusions this study successfully designed and developed the "youth innovation village" system to support the digital transformation of karang taruna mekarjati. the system was built using the laravel framework, php, mysql, and bpmn-based process modeling, with an iterative development approach following the agile scrum methodology. the research addressed the inefficiencies of manual administrative processes within karang taruna, particularly in membership registration, event management, leadership training, and msme/startup monitoring. by implementing this digital system, the organization experienced significant improvements in operational efficiency, data accuracy, and member engagement. the system's functional testing, performance evaluation, and user feedback analysis demonstrated its effectiveness. results from unit testing, integration testing, and user acceptance testing (uat) showed that the core functionalities, such as membership registration and event management, performed smoothly, while minor improvements were needed in startup monitoring filters and ui accessibility. additionally, performance testing indicated that the system efficiently handled concurrent users and real-time data processing, ensuring reliability for organizational use. user feedback highlighted high satisfaction levels, with 92% of respondents finding the system easy to use, 89% satisfied with its responsiveness, and 91% expressing overall approval. administrators particularly benefited from real-time analytics and centralized data management, which streamlined decision-making processes. however, some users suggested enhancements, including a mobile application for better accessibility and expanded financial management features to support karang taruna's economic initiatives. comparing the manual system with the digital system, the findings confirmed that the new platform significantly improved workflow efficiency, data security, and organizational communication. the centralized dashboard, automated membership tracking, and integrated event management replaced manual record-keeping and fragmented communication, reducing errors and delays. this aligns with previous studies emphasizing digital transformation in non-profit organizations, reinforcing that technology adoption enhances transparency, engagement, and operational effectiveness. despite the positive outcomes, challenges remain, such as user adaptation for less tech-savvy members, dependency on stable internet connections, and the need for further system enhancements. future research and development should focus on: 1. developing a mobile application to improve accessibility and engagement. 2. implementing financial tracking features to support budget planning and fundraising initiatives. 3. leveraging ai and data analytics to predict event participation and optimize resource allocation. vol. 6, no.2, july 2025 | 90 acknowledgements the authors would like to express their sincere gratitude to karang taruna mekarjati for their cooperation and valuable insights throughout the research and development of the "youth innovation village" system. their active participation and feedback played a crucial role in shaping the system to meet the organization's needs effectively. we also extend our appreciation to horizon university for providing academic guidance and resources that supported this research. special thanks to the faculty of informatics and computer technology (fict) for their encouragement and constructive input during the study. furthermore, we acknowledge the contributions of all participants, volunteers, and survey respondents who provided essential feedback and suggestions for system improvements. their involvement has been instrumental in refining the usability and functionality of the platform. lastly, we would like to thank our mentors, colleagues, and family members for their continuous support, motivation, and encouragement throughout the research and development process. this project would not have been possible without their unwavering guidance and assistance. references [1]. chigunta, f. j. (2022). youth entrepreneurship: meeting the key policy challenges. education development center. https://search.worldcat.org/title/youth-entrepreneurship-%3a-meetingthe-key-policy-challenges/oclc/51902859 [2]. drucker, p. f. (n.d.). innovation and entrepreneurship. harper business. https://www.harpercollins.com/products/innovation-and-entrepreneurship-peter-fdrucker?variant=32118080077858 [3]. highsmith, j. (2010). agile project management: creating innovative products, 2nd edition. addison-wesley professional. https://www.informit.com/store/agile-project-managementcreating-innovative-products-9780321658395 [4]. indonesian ministry of social affairs. (n.d.). karang taruna: strengthening social welfare through youth empowerment. indonesian ministry of social affairs. https://www.kemsos.go.id/uploads/topics/16327335291494.pdf [5]. khalili, b., & smyth, a. w. (2024). sod-yolov8—enhancing yolov8 for small object detection in aerial imagery and traffic scenes. sensors, 24(19), 6209. https://doi.org/10.3390/s24196209 [6]. martin fowler. (n.d.). uml distilled: a brief guide to the standard object modeling language, 3rd edition. addison-wesley professional. https://www.informit.com/store/uml-distilled-abrief-guide-to-the-standard-object-modeling-9780321193681 [7]. mcnutt, j., guo, c., goldkind, l., & an, s. (2018). technology in nonprofit organizations and voluntary action. brill. https://doi.org/10.1163/9789004378124 [8]. nabila ahmed nikita-, k. s. a., azher uddin shayed, mir abrar hossain, & obyed ullah khan. (2024). digital transformation in non-profit organizations: strategies, challenges, and successes. advanced international journal of multidisciplinary research, 2(5), 1097. https://doi.org/10.62127/aijmr.2024.v02i05.1097 [9]. sommerville, i. (2021). engineering software products: an introduction to modern software engineering, 1st edition. engineering/p200000003243/9780137524846?tab=title-overview [10]. the 2020 scrum guide. (n.d.). https://scrumguides.org/scrum-guide.html [11]. youth global programme for sustainable development & peace. (n.d.). [12]. priyatna, b. (2019). penerapan metode user centered design (ucd) pada sistem pemesanan menu kuliner nusantara berbasis mobile android. jurnal accounting information system (aims), 2(1), 17-30. [13]. priyatna, b., hananto, a. l., & nova, m. (2020). application of uat (user acceptance test) evaluation model in minggon e-meeting software development. systematics, 2(3), 110-117. [14]. ismail, d. a., huda, b., hilabi, s. s., & priyatna, b. (2024). penerapan desain ui/ux pada sistem penjualan berbasis web dengan metode desain thingking. innovative: journal of social science research, 4(2), 5737-5748. https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht https://www.zotero.org/google-docs/?cooyht vol. 6, no.2, july 2025 | 91 [15]. priyatna, b., hilabi, s. s., heryana, n., & solehudin, a. (2019). aplikasi pengenalan tarian dan lagu tradisional indonesia berbasis multimedia. systematics, 1(2), 89-98. vol.6, no.2, july 2025 | 92 p-issn: 2715-2448 | e-issn: 2715-7199 vol.6 no.2 july 2025 buana information technology and computer sciences (bit and cs) detection of hijaiyah letters handwritten in early childhood using yolo v8 hidayatul mustagfiroh1, aviv yuniar rahman2, rangga pahlevi putra3 1,2,3 department of informatic engineering, universitas widya gama, malang, indonesia e-mail: hidayatulm356@gmail.com1, aviv@widyagama.ac.id2, rangga@widyagama.ac.id3 received: 2025/05/27 | revised: 2025/06/26 | accepted: 2025/07/28 abstract this study investigates the effectiveness of the yolov8 (you only look once version 8) algorithm in detecting handwritten hijaiyah letters among early childhood learners. the introduction of technology in early childhood education is essential for enhancing literacy skills, particularly in learning the arabic alphabet, which is crucial for reading the quran. this research addresses the challenges faced by educators in assessing children's handwriting, which often lacks consistency and objectivity. a dataset of 3,780 images of handwritten hijaiyah letters was collected from children at ra baipas roudlotul jannah, including various writing styles to ensure the model's robustness. prior to training, the images underwent preprocessing steps such as resizing, normalization, and data augmentation techniques like rotation and flipping to enhance the quality and diversity of the training data. the yolov8 model was trained using an 80-10-10 split for training, validation, and testing datasets. evaluation metrics such as precision, recall, and mean average precision (map) were used. the results showed that yolov8 achieved an impressive accuracy of 96.08% in detecting handwritten hijaiyah letters, with high precision and recall rates further validating the model's reliability. this research highlights the potential of integrating advanced object detection algorithms like yolov8 into educational practices. by providing real-time feedback, the system can significantly enhance the learning experience for young children, facilitating their understanding and mastery of the arabic alphabet. future research should focus on expanding the dataset and refining the model to address handwriting variability challenges and improve accuracy. keywords: arabic alphabet, early childhood education, handwriting recognition, hijaiyah letters, object detection, and yolo. i. introduction the hijaiyah letter is a letter to arrange 28 words in arabic with different forms [1]. however, there are other sources that mention it with other numbers. among them, there are 28 and 30 [2]. including lamalif, which some people consider to be a different letter [3]. so that without lamalif the total number is 29 hijaiyah letters. in addition, the generally accepted count, excludes lamalif and combines alif with hamzah [4]. so some sources dispute this calculation, emphasizing the complexity of the arabic script [5] . the letter begins with alif and ends with ya. hijaiyah letters are also an integral part of the quran, both as a basis for reading it and understanding its contents [6]. therefore, learning, memorizing, and understanding hijaiyah letters is the first stage to be able to read and understand the quran. however, not everyone can immediately know and understand the basic science of the quran, namely hijaiyah letters, especially early childhood in kindergarten, namely ra baipas roudlotul jannah. they generally do not know the shape, how to read, and how to write all the hijaiyah letters properly and correctly. understanding the cognitive development stage of children is very important to adjust the teaching method. the learning process involves recognizing letter shapes, understanding simple words, and reading the signs of the letters [7]. vol.6, no.2, july 2025 | 93 then there are obstacles when what is assessed is the results of children's handwriting work, for example on the object of handwriting hijaiyah letters of the work / work of early childhood. subjectively humans are able to provide an assessment of the work but there are times when it is less consistent and difficult to determine with certainty the level of similarity of hijaiyah letter handwriting to the hijaiyah letters used as a reference [4]. however, while automated systems can improve accuracy, they may not fully capture the nuances of individual handwriting styles, which remain important in evaluating children's learning progress. one solution to maximize the learning process and useful for teachers to make it easier to correct the handwriting of the hijaiyah letters of their students is to use object detection [8]. this research tries to develop object detection of hijaiyah letter writing in early childhood, one of which is the application of object detection, especially using the yolo (you only look once) method, can significantly improve the learning process to recognize hijaiyah letters in early childhood education [9],[15]. the detection system using yolo is proven to be faster and more accurate to detect an object in an image or image so that it is most suitable if applied to the case taken by the researcher [10]. using object detection can help teachers in recognizing and distinguishing each of the 30 hijaiyah letters that a person will learn. ii. methods this chapter will discuss the methods used in the process of detecting children's handwritten hijaiyah letters. this research method is a method used to detect early childhood handwritten hijaiyah letter objects using yolov8. this research uses a quantitative approach with an experimental design that aims to test the effectiveness of the hijaiyah letter detection system. in this research design stage, it is presented in the flowchart illustration in figure 1. the steps are carried out sequentially in order to get maximum results in writing the final report. . figure 1. research design. (source: personal preparation) vol.6, no.2, july 2025 | 94 the image above illustrates the step-by-step process involved in building and training a model to recognize handwritten hijaiyah letters. 1. dataset retrieval the first step is the retrieval of datasets, which involves taking photographs of early childhood handwritten hijaiyah letters using a printer scanner. the dataset, consisting of 3,780 images, is categorized into 30 classes, with each class containing 126 images of different hijaiyah letters. these images serve as the training samples for the system being built. 2. data preprocessing once the dataset is collected, it goes through a preprocessing stage. during this stage, the images are resized to ensure consistency in dimensions across the entire dataset. this step is critical for optimizing the images and ensuring that they meet the input requirements of the model being trained. consistent image size allows the model to process and learn from the data efficiently. 3. labeling dataset the next step is labeling the dataset, which involves marking each image with a bounding box using labeling software. this step is essential for training the model to recognize the objects (in this case, hijaiyah letters) within the images. each bounding box represents the object in the image, allowing the system to identify and classify it accurately. 4. split dataset after labeling the dataset, the data is divided into three parts: training data, validation data, and testing data. the training data is used to train the model, while the validation data is utilized to evaluate the model's performance during training. the testing data, which has never been seen by the model before, is used to assess the model's ability to generalize and perform accurately on new, unseen data. 5. training at this stage, the divided data—training and testing—are accessed through an api in google colab. the model is trained using the yolov8 algorithm, with accuracy being the primary parameter for evaluating performance. if the initial test results show low accuracy, retraining is performed. this retraining aims to optimize the model to improve its accuracy and overall prediction capability. 6. detection result the detection process begins by feeding the collected images into the system, which uses the yolov8 algorithm to analyze the images and recognize objects. the model then marks the detected objects with bounding boxes, providing a clear distinction between the objects and the background. the output of this process is an accuracy score, which reflects how well the model is able to identify the objects within the images. 7. model evaluation finally, the model's performance is evaluated based on its ability to recognize hijaiyah letters in previously unseen images. evaluation metrics include the detection accuracy, inference time, and other relevant parameters that measure the effectiveness of the model. this evaluation ensures that the model performs well in recognizing objects under real-world conditions. vol.6, no.2, july 2025 | 95 iii. results and discussions the results of this study explain the object detection of early childhood handwritten hijaiyah letters using yolov8. with this discussion, it can be seen the success in detecting children's handwriting objects. datasets that have been labeled will be resized and divided into 3 types of data, namely 80% train data, 10% valid data and 10% test data. to facilitate the author in dividing label data quickly without having to sort out one by one [12]. the dataset sharing process is shown in figure 2. figure 2. split dataset anaconda (source: personal preparation) after the installation of the yolov8 algorithm system, the training process with the yolov8 algorithm uses the model training configuration, namely images with a width resolution of 640 with adjusting height, a total of 50 epochs and 10 batches. the model configuration process can be seen in figure 3. figure 3. training model yolov8 (source: personal preparation) the results of the evaluation of the yolov8 model on the validation set consisting of 378 images showed excellent performance, with an accuracy of 0.973 and a recall of 0.5597. map50 reached 0.6129, and map 50-95 of 0.4147. the average processing time per image is 4.5 ms for inference, which indicates its reliable real-time detection capabilities. the results of the evaluation of the yolov8 model can be seen in figure 4. figure 4. yolov8 evaluation results (source: personal preparation) after the model will be evaluated using test data to measure its overall performance. this involves using evaluation metrics such as precision, recall, and map to measure how well the model can detect hijaiyah datasets in the image [14]. results as shown in table 1. vol.6, no.2, july 2025 | 96 table 1. results of the evaluation of the yolov8 model (source: personal preparation) class name box(p) recall map 50 map5095 class name box(p) recall map 50 map50-95 alif 0,947 0,949 0,961 0,766 tha 0,972 1 0,995 0,867 ba' 0,956 0,789 0,943 0,79 dha 0,973 1 0,995 0,783 ta 0,925 1 0,938 0,723 ain 0,997 1 0,995 0,686 tsa 0,92 0,828 0,862 0,622 ghain 0,981 0,938 0,946 0,698 jim 0,984 1 0,995 0,698 fa 0,996 1 0,995 0,778 kha 0,777 1 0,871 0,685 qaf 0,979 1 0,979 0,649 kho 0,99 1 0,995 0,773 kaf 0,92 1 0,979 0,641 dal 0,848 1 0,942 0,705 lam 0,937 0,931 0,934 0,602 dzal 0,764 0,778 0,827 0,616 mim 0,986 1 0,968 0,687 ra 0,981 1 0,938 0,683 nun 0,925 0,867 0,914 0,674 zai 0,939 0,853 0,949 0,742 waw 0,862 0,965 0,968 0,617 sin 0,993 1 0,995 0,78 ha 1 0,964 0,995 0,761 syin 0,981 1 0,995 0,731 lam alif 0,821 1 0,898 0,561 shad 0,974 1 0,995 0,7 hamzah 1 0,925 0,995 0,75 dhad 0,958 0,923 0,95 0,669 ya 0,922 0,984 0,983 0,786 the results of the evaluation of the yolov8 model on the validation set consisting of 378 images showed excellent performance, with an accuracy of 0.973 and a recall of 0.5597. map50 reached 0.6129, and map 50-95 reached 0.4147. the average processing time per image is 4.5 ms for inference, which shows reliable real-time detection capabilities. figure 5. precision-recall curve (source: personal preparation) from figure 5. it can be seen that the level of accuracy is quite good, even until someone touches the number 16 which means very accurate, for example the letter ghain. the detection of hijaiyah letters using yolov8 went well and the accuracy value was quite high. in table 2. explain the results of the hijaiyah letter detection test vol.6, no.2, july 2025 | 97 table 2. test results (source: personal preparation) test data dataset class detection results level of certainty shad 1.0 tsa 0.8 dzal 0.8 alif 1.0 ba 0.9 iv. conclusions the application process for the detection of hijaiyah letters written by early childhood used yolov8 with a lancer and succeeded in accurately detecting the presence of breeders. the dataset used consisted of 3790 images on hijaiyah letters which were divided into 30 classes. this dataset is divided into 3 parts: training data (80%), validation (10%), and testing (10%) [12]. detection using the yolo model yielded an accurate accuracy of 0.9608 drawn from the map50 results, 0.973 precision, and 0.98 recalls. the graph shows a precision average value (map) of 0.962 for all classes. this presentation shows that the real-time object detection system using yolov8 provides accurate and reliable results when tested. vol.6, no.2, july 2025 | 98 references [1] y. mohamed, s. ismail, and y. suryadama, “pronunciation of hijaiyyah’s letter for new quranic learners a contrastive analysis study,” ulum islam., vol. 36, no. 01, pp. 73–82, 2024, doi: 10.33102/uij.vol36no01.553. [2] asmaa rafat elsaied, “relationship between numbers and letters,” j. math. syst. sci., vol. 6, no. 8, pp. 335–337, 2016, doi: 10.17265/2159-5291/2016.08.005. [3] l. sarifah, s. khotijah, and m. k. khaliqah, “identification of hijaiyah letters image using extreme learning machine method,” j. mat. stat. dan komputasi, vol. 20, no. 1, pp. 90–101, 2023, doi: 10.20956/j.v20i1.27158. [4] m. m. bahjat, e. sayed, m. salem, and a. a. ghafoor, “a lexicon of basic vocabulary in the holy quran: the ‘hamza’ character as a model,” vol. 6, no. 1, pp. 56–67, 2023, doi: 10.18860 /ijazarabi.v6i1.17642. [5] a. saber, a. taha, and k. abd el salam, “a comprehensive approach to arabic handwriting recognition: deep convolutional networks and bidirectional recurrent models for arabic scripts,” int. j. telecommun., vol. 04, no. 02, pp. 1–11, 2024, doi: 10.21608/ijt.2024.291347.1052. [6] r. f. rahmat, f. akbar, m. f. syahputra, m. a. budiman, and a. hizriadi, “an interactive augmented reality implementation of hijaiyah alphabet for children education,” j. phys. conf. ser., vol. 978, no. 1, 2018, doi: 10.1088/1742-6596/978/1/012102. [7] k. stein-smith et al., “the independent self-directed language learner and the role of the language educator — expanding access and opportunity kathleen,” j. lang. teach. res., vol. 14, no. 2, pp. 5–13, 2023, doi: https://doi.org/10.17507/jltr.1402.01. [8] siti mahrami ivlatia, nina wandana, dita andini harahap, aslam annashir, and sahkholid nasution, “analisis kompetensi penulisan huruf hijāiyah tunggal pada siswa mis ummi lubuk pakam,” semant. j. ris. ilmu pendidikan, bhs. dan budaya, vol. 2, no. 1, pp. 188–200, jan. 2024, doi: 10.61132/semantik.v2i1.284. [9] h. sidi, a. yuniar, and f. marisa, “expression detection of children with special needs using yolov4-tiny,” vol. 16, no. 3, pp. 221–227, 2025. [10] a. y. rahman and z. zakaria, “hybrid yolov8 and fast r-cnn for accurate schematic detection in power distribution networks,” ieee access, vol. pp, p. 1, 2025, doi: 10.1109/access.2025.3561279. [11] j. j. p. jansen, c. heavey, t. j. m. mom, z. simsek, and s. a. zahra, “scaling-up: building, leading and sustaining rapid growth over time,” j. manag. stud., vol. 60, no. 3, pp. 581– 604, may 2023, doi: 10.1111/joms.12910. [12] v. c. raykar and a. saha, “data split strategies for evolving predictive models,” lect. notes comput. sci. (including subser. lect. notes artif. intell. lect. notes bioinformatics), vol. 9284, pp. 3–19, 2015, doi: 10.1007/978-3-319-23528-8_1. [13] z. q. zhao, p. zheng, s. t. xu, and x. wu, “object detection with deep learning: a review,” ieee trans. neural networks learn. syst., vol. 30, no. 11, pp. 3212–3232, 2019, doi: 10.1109/tnnls.2018.2876865. vol.6, no.2, july 2025 | 99 [14] r. padilla, w. l. passos, t. l. b. dias, s. l. netto, and e. a. b. da silva, “a comparative analysis of object detection metrics with a companion open-source toolkit,” electron., vol. 10, no. 3, pp. 1–28, feb. 2021, doi: 10.3390/electronics10030279. [15] priyatna, b., rahman, t. k. a., hananto, a. l., hananto, a., & rahman, a. y. (2024). mobilenet backbone based approach for quality classification of straw mushrooms (volvariella volvacea) using convolutional neural networks (cnn). joiv: international journal on informatics visualization, 8(3-2), 1749-1754. vol. 5, no.2 june 2024 | 51 apriori algorithm and market basket analysis to uncover consumer buying patterns: case of a kenyan supermarket edwin omol1*, dorcas onyango2, lucy mburu3, paul abuonji4 1,2,3,4 department of computing and information technology, kenya highlands university p. o. box 123 20200 kericho, kenya 1* omoledwin@gmail.com, 2dawino2011@gmail.com, 3mburul@kcau.ac.ke, 4pabuonji@kcau.ac.ke abstract: this article presents a study on utilizing the apriori algorithm and market basket analysis (mba) to reveal consumer buying patterns in supermarkets. the aim of this research is to explore the effectiveness of these data mining techniques in revealing valuable insights that can inform marketing strategies and enhance the overall shopping experience for customers. this study centered on improving customer loyalty within the supermarket setting through the utilization of cutting-edge information technology and programming applications, including python. specifically, the apriori algorithm libraries of the python language were employed to identify frequent item sets and derive 42 association rules, which shed light on product affinities and co-purchasing patterns. by deriving association rules from the frequent item sets, the study identified the significance of strategically placing frequently purchased products to enhance revenue generation. in conclusion, the application of the apriori algorithm and market basket analysis in this case of a kenyan supermarket has proven to be a valuable approach for uncovering consumer buying patterns, providing a competitive edge in the dynamic retail industry. keywords: market basket analysis, consumer buying patterns, data mining techniques, marketing strategies i. introduction: in the highly competitive retail industry driven by digital technologies [1], understanding consumer behavior and buying patterns is crucial for supermarkets to tailor their marketing strategies and enhance customer satisfaction [2]. with the vast amount of transactional data generated at supermarkets, technological innovations [6] like data mining techniques, particularly market basket analysis, have emerged as powerful tools to gain valuable insights into consumer purchase behavior. this article aims to investigate the buying patterns of consumers at a prominent kenyan supermarket using market basket analysis [3-5]. market basket analysis involves analyzing customers' purchase transactions to identify associations between products frequently bought together. by examining these patterns, supermarkets can optimize product placements, offer personalized promotions, and improve inventory management [3]. the insights derived from this analysis can help supermarkets enhance their overall shopping experience, increase customer loyalty, and boost profitability [5,7]. this study delves into the purchasing patterns of society stores’ diverse consumer base. society stores, a rapidly expanding kenyan supermarket catering to the mass market, places its primary emphasis on providing superior products at budget-friendly rates. it has established outlets in various kenyan locations, including thika, naivasha, ruiru, maua, limuru, meru, and mombasa, kenya's bustling p-issn: 2715-2448 | e-issn : 2715-7199 vol.5 no.2 june 2024 buana information technology and computer sciences (bit and cs) mailto:omoledwin@gmail.com mailto:dawino2011@gmail.com mailto:mburul@kcau.ac.ke mailto:pabuonji@kcau.ac.ke vol. 5, no.2 june 2024 | 52 economic hubs. by examining the association rules between different products, we aimed to identify popular product combinations and uncover consumer preferences. additionally, we explored the influence of demographics, such as age, gender, and income [10-12], on purchase behavior to gain a comprehensive understanding of the factors driving consumer choices. ii. literature review: the study of consumer behavior has always been a crucial aspect of marketing and retail management. understanding the preferences, buying habits, and patterns of consumers is essential for businesses to tailor their marketing strategies, optimize product placements, and enhance customer satisfaction [8,2]. over the years, various analytical techniques have been developed to analyze consumer purchasing behavior, and one such powerful method is the apriori algorithm coupled with market basket analysis (mba) [3,4]. the apriori algorithm is a widely used association rule mining technique in data mining and machine learning. it aims to discover interesting relationships or associations between items in large datasets, particularly in transactional databases. the algorithm is highly efficient and effective in identifying frequent item sets, which are groups of items that appear together frequently in transactions. by using the apriori algorithm, researchers and marketers can extract valuable association rules that reveal hidden patterns in consumer shopping habits [13]. market basket analysis, on the other hand, is a practical application of the apriori algorithm in retail and e-commerce industries. it involves the analysis of customer transactions to identify the co-occurrence of products that tend to be purchased together [14]. this analysis provides valuable insights into crossselling opportunities and enables businesses to design effective promotional strategies and optimize store layouts. in the context of a kenyan supermarket, where consumer behavior may be influenced by cultural, social, and economic factors unique to the region [11], the combination of the apriori algorithm and market basket analysis presents an excellent opportunity to uncover meaningful patterns in consumer buying behavior. understanding which products are frequently purchased together can help the supermarket enhance product bundling, offer personalized recommendations, and optimize inventory management. several studies have successfully applied the apriori algorithm and market basket analysis to investigate consumer buying patterns in various retail settings worldwide. similar research has been conducted in supermarkets, grocery stores, online shopping platforms, and other retail environments to explore consumer preferences and optimize business strategies. for instance, in a study conducted by xie, [22] in a chinese supermarket, the apriori algorithm was employed to analyze transactional data and identify significant association rules. the findings revealed interesting patterns in consumer shopping behavior, leading to improved store layouts and targeted marketing campaigns. additionally, a study by ünvan, [20] applied market basket analysis to e-commerce data in the united states, shedding light on product affinities and uncovering opportunities for cross-selling and upselling. the comparative study by chen & zhang, [4] evaluated the performance of the apriori and fpgrowth algorithms in market basket analysis. chen and zhang analyze their efficiency, scalability, and ability to reveal consumer buying patterns, shedding light on the strengths and weaknesses of each method [4]. li and tan conducted a review of sequential pattern mining techniques in market basket analysis. the study discussed the limitations of traditional association rule mining and highlights the importance of considering the temporal order of transactions to capture consumer buying patterns effectively [9]. in their survey paper, fournier-viger et al. [18] present an in-depth analysis of apriori-based algorithms for frequent itemset mining, including their applications in uncovering consumer buying patterns. the research discussed various modifications and improvements to the apriori algorithm and vol. 5, no.2 june 2024 | 53 their impact on market basket analysis. zhang and liu [19] proposed an improved version of the apriori algorithm to mine consumer buying patterns. the study demonstrated how this modification enhances the efficiency and effectiveness of market basket analysis, enabling businesses to gain valuable insights into consumer preferences and behavior. smith & johnson, [14] study explored the application of the apriori algorithm in retail settings to identify consumer buying patterns. the research demonstrated how the algorithm efficiently generates frequent item sets and association rules, providing valuable insights into consumer behavior in the retail industry while wang and lee investigated the utilization of market basket analysis and big data techniques to uncover consumer buying patterns in e-commerce. the study highlighted the advantages of using these data mining techniques in the digital retail context to enhance marketing strategies and customer experience [21]. according to sornalakshmi et al. [17], the utilization of the apriori algorithm in market basket analysis offers several benefits. first, it efficiently generates frequent item sets through the elimination of infrequent item sets, resulting in reduced computational complexity. additionally, it effectively derives association rules from frequent item sets, allowing businesses to discern significant relationships among items. furthermore, the apriori algorithm is widely embraced in retail for market basket analysis and has proven its effectiveness in revealing purchasing patterns. its simple approach and intuitive nature also facilitate relatively straightforward implementation in various programming languages. moreover, the algorithm's scalability permits its application to vast transactional databases, making it well-suited for analyzing extensive retail datasets [17]. the apriori algorithm's generation of candidate item sets results in a combinatorial explosion, leading to high memory and computational requirements [22]. this limitation poses challenges when analyzing large transactional databases, causing scalability issues [18]. researchers have noted that multiple passes over the data may be necessary, making it less efficient for big data analysis [14]. given the growing interest in understanding consumer behavior and the increasing availability of large-scale transactional data, the use of the apriori algorithm and market basket analysis in supermarkets and retail industries has become increasingly relevant and valuable [20]. in summary, the combination of the apriori algorithm and market basket analysis offers a robust approach to unearthing meaningful insights into consumer buying patterns. by applying this methodology to a kenyan supermarket, we aim to contribute to the body of knowledge on consumer behavior in the region and provide actionable recommendations for the supermarket's marketing and operational strategies. iii. method the methodology employed involved a multi-step process combining data collection, data preprocessing, and the application of the apriori algorithm in market basket analysis. see fig. 1. vol. 5, no.2 june 2024 | 54 fig. 1: study method 1. data collection the first step in this study was the collection of transactional data from the society stores supermarket. the data included detailed information about individual customer transactions, such as the products purchased, the transaction date, and the transaction amount. the data was obtained with the permission and cooperation of the supermarket management to ensure data privacy and confidentiality [10]. 2. data preprocessing once the data was collected, it underwent thorough preprocessing to ensure its quality and readiness for analysis. data preprocessing involved tasks such as data cleaning, handling missing values, and transforming categorical variables into a suitable format for market basket analysis. additionally, any irrelevant or redundant data was removed to focus solely on transactional information relevant to the study. 3. market basket analysis (mba) the core of this study's methodology lies in the application of market basket analysis. mba was performed on the preprocessed transactional data to identify frequent item sets and uncover hidden patterns of product associations. association rules were generated to reveal the likelihood of customers purchasing specific products together. the analysis utilized established algorithms like the apriori algorithm [4] to efficiently mine association rules from the transactional data. 4. frequent item sets and association rules in order to explore frequent item sets within the dataset denoted as "my_basket_sets," employing the apriori algorithm was necessary. a minimum support threshold of 0.01, equivalent to 1% of the total transactions, was applied. the output of this analysis showcased the frequent item sets, along with relevant metrics such as support, confidence, and lift values, among others, which facilitated the establishment of association rules. this market basket analysis aided in identifying market implications aligned with consumer preferences, grouping products based on buying habits, and streamlining the search process. the apriori algorithm was commonly employed for this purpose, as it effectively uncovered combinations of products that are frequently purchased together. •data collection data aquisition •data preparation •feature selection data processing •apriori algorithm & market basket analysis •cluster analysis •interpretation and insights data visualization vol. 5, no.2 june 2024 | 55 5. interpretation and insights the results obtained from market basket analysis and cluster analysis were thoroughly interpreted to extract meaningful insights into consumer purchase behavior at the society stores supermarket. the association rules highlighted which products are frequently purchased together, indicating potential crossselling opportunities and product bundling strategies. the clustering results provided a deeper understanding of different customer segments and their unique buying preferences. iv. results and discussions the study employed the apriori algorithm and market basket analysis (mba) to examine consumer buying patterns in a kenyan supermarket. the analysis was based on transactional data collected over a specific period, capturing the purchases of various products by individual customers. the objective was to identify frequent itemsets and association rules that could shed light on consumer preferences and uncover meaningful patterns in their shopping behavior. 1. identification of frequent itemsets fig. 2: top 20 items purchased by customers vol. 5, no.2 june 2024 | 56 fig. 3: item co-occurrence fig. 4: frequent item-sets fig. 2 illustrates the successful application of the apriori algorithm, which effectively identified 20 frequent item sets commonly purchased by the majority of consumers. the analysis revealed that coffee, cake, bread, tea, and pastry were the top five items frequently found in shopping baskets, indicating that a significant number of consumers are coffee and tea enthusiasts who often accompany their beverages with bread, cake, pastry, and sandwiches. vol. 5, no.2 june 2024 | 57 furthermore, figure 2 demonstrated the identification of relationships between items that tend to cooccur frequently in transactions. the analysis unveiled sets of products showing strong co-occurrence patterns, suggesting that consumers tend to buy these items together during their shopping trips, as illustrated in fig. 3. the results in figure 3 particularly highlighted the co-occurrence of coffee, bread, and cake in numerous shopping baskets, potentially attributed to their complementary nature. in addition, the analysis, as depicted in fig. 4, revealed which products were commonly purchased together by customers, along with their respective support levels. the output highlighted coffee as a significantly prevalent item in the majority of shopping baskets, indicating its popularity among consumers. 2. association rules analysis fig 5: product association rules through the application of association rule mining to the frequent item sets, the study extracted meaningful and actionable 42 rules as displayed in fig. 5. these rules shed light on the likelihood of customers purchasing specific items based on their previous purchases. the data-derived association rules indicate that coffee appears in approximately 47% of all baskets when observing a customer's behavior. conversely, the occurrence of toast intake is observed at a rate of 3%. additionally, it was observed that vol. 5, no.2 june 2024 | 58 70% of customers who purchase toast also buy coffee, indicating a strong preference for coffee over toast due to a high confidence level and a lift metric of 1.47. furthermore, the association rules suggest a probability of 1% for spanish brunch purchases. the concurrent support level of coffee with spanish brunch is measured at 1%, and these items exhibit a confidence level of 59%. this implies that spanish brunch serves as a secondary preference for most consumers compared to coffee. regarding medialuna, the association rules indicate a probability of 6% for its purchases. the support level of coffee with medialuna is measured at 3%, with 56% of those who buy medialuna also purchasing coffee. this suggests a significant association between coffee and medialuna, indicating that they are often bought together. in contrast, the association rules derived from the data show a 4.9% likelihood of encountering both coffee and tea in customers' purchases. meanwhile, the occurrence of cake in conjunction with both coffee and tea is observed at a rate of 10%. furthermore, 9.6% of customers who purchase cake also buy both coffee and tea, suggesting that consumers tend to opt for either tea or coffee, but not both. 3. product affinities and cross-selling opportunities fig 6: item affinities vol. 5, no.2 june 2024 | 59 the examination yielded significant product affinities, indicating a pattern of frequent co-purchases by customers. as depicted in fig. 6, the items strongly associated with coffee purchases include toast, medialuna, spanish brunch, pastry, and alfajores. this affinity could be attributed to consumer preferences. however, certain items, such as cake and bread, displayed lower affinities with coffee. this could be explained by the perception among consumers that cakes and bread contain higher sugar content, leading to reduced consumption in conjunction with coffee. 4. time period trends and purchase behavior fig 7: time period trends the research delved into the seasonal trends in consumer purchasing behavior, with a focus on different time periods. transactional data analysis allowed the study to identify changes in buying patterns throughout the day—morning, evening, afternoon, and night. fig. 7 displays the top 10 items commonly ordered by consumers during these specific time frames. during morning and afternoon hours, consumers showed a higher tendency to purchase significant quantities of coffee and bread. however, no coffee was observed in night-time orders, where vegan feast and hot chocolate toppings were more prevalent in most shopping baskets. this suggests that individuals generally prefer coffee during morning and afternoon hours to stay refreshed and productive throughout the day. conversely, during nighttime, juice, and mineral water were the preferred choices. it could be possible that consumers opt for hydrating options in the evening and may also purchase sweet treats for their families during this period. 5. optimizing store layout the analysis of consumer buying patterns provided insights into the optimization of the supermarket's store layout. by strategically placing frequently co-purchased items closer together, the supermarket can create a more convenient shopping experience for customers and potentially increase impulse purchases. in conclusion, the application of the apriori algorithm and market basket analysis proved highly valuable in uncovering consumer buying patterns in the kenyan supermarket. the findings provided actionable insights for the supermarket to optimize marketing strategies, enhance product bundling and cross-selling opportunities, and improve customer satisfaction. by leveraging these insights, the vol. 5, no.2 june 2024 | 60 supermarket can stay competitive in the market and provide a more personalized and enjoyable shopping experience for its customers. the findings of this study demonstrate the effectiveness of employing the apriori algorithm and market basket analysis (mba) to uncover valuable insights into consumer buying patterns in a kenyan supermarket. the analysis of transactional data provided significant results that can be utilized by the supermarket to enhance its marketing strategies, optimize product placements, and improve customer satisfaction. 1. market basket analysis reveals frequent item sets the application of the apriori algorithm successfully identified frequent item sets, representing sets of products that are frequently purchased together by customers. by promoting cake or pastry discounts alongside coffee or tea purchases can stimulate additional sales and create a sense of convenience for consumers looking for complementary treats. according to ünvan [20], it is recommended to position related products in close proximity to one another. moreover, given the popularity of coffee, bread, and cake as co-occurring items, restaurants, cafes, and supermarkets can strategically design their menus and displays to highlight these combinations. creating visually appealing displays showcasing these items together can influence customer choices and drive impulse purchases. additionally, in leveraging the information about coffee's high prevalence in shopping baskets, businesses can design loyalty programs focused on coffee enthusiasts. offering exclusive benefits or rewards for coffee-related purchases can incentivize repeat visits and build customer loyalty. 2. association rules offer actionable insights by deriving association rules from the frequent item sets, the study identified marketing implications, that could enhance customer satisfaction by catering to their preferences and needs. the association rule indicating that 70% of customers who purchase toast also buy coffee suggests a strong preference for coffee over toast. this implies that promoting coffee in combination with toast or as a complementary item could further boost coffee sales. these findings align with the research conducted by suryadi and islami [18], which highlights the significance of strategically placing frequently purchased products to enhance revenue generation. while spanish brunch has a low occurrence rate (1%), it is the second preference for many customers after coffee. to attract more customers interested in spanish brunch, targeted marketing campaigns or promotions highlighting this item could be implemented. medialuna is chosen by 6% of customers, and 56% of those who buy medialuna also buy coffee. promoting medialuna alongside coffee or creating special offers for this combination could enhance sales and encourage customers to try both items together. the association rule indicating that 9.6% of customers who purchase cake also buy coffee and tea suggests that consumers tend to choose either coffee or tea, but not both. this insight can be utilized to offer specific deals or promotions that encourage customers to pair cake with their preferred hot beverage (coffee or tea). finally, tea appears in approximately 4.9% of customer baskets, which is relatively low compared to coffee. marketing efforts could be directed towards promoting tea to increase its occurrence in customer purchases and potentially expand its customer base. 3. cross-selling and revenue generation the study revealed strong product affinities and co-purchasing patterns, which offer opportunities for cross-selling. given the strong product affinities between coffee and items such as toast, medialuna, spanish brunch, pastry, and alfajores, there is an opportunity for the business to create bundle offers or cross-selling promotions. by strategically pairing these items with coffee, the business can encourage customers to make additional purchases and potentially increase their average transaction value. the findings presented are consistent with the observations made by hermina, aishwaryalakshmi, and vol. 5, no.2 june 2024 | 61 gopalakrishnan [5], who suggest that there is a possibility of mineral water being frequently purchased alongside other products, thereby offering opportunities for strategic product placement and cross-selling strategies. furthermore, the products that are frequently bought together with coffee can be strategically placed near the coffee counter. this can influence impulse buying and encourage customers to add these complementary items to their coffee orders. understanding that some items, like cakes and bread, have lower affinities with coffee due to perceived higher sugar content, the business can promote healthier alternatives or low-sugar options for health-conscious customers. this could involve introducing sugar-free or reduced-sugar variations of cakes and bread on the menu. 4. understanding time period trends the study also highlighted seasonal trends in consumer purchasing behavior. offering a variety of coffee options and bread choices during morning and afternoon hours can attract more customers during those periods. additionally, featuring vegan feast and hot chocolate toppings in the evenings can appeal to consumers seeking comfort or indulgence during nighttime. the store could also offer promotional deals on coffee at night or bundling sweet treats with juice purchases can encourage consumers to make purchases during off-peak hours. moreover, the store could ensure a swift and efficient coffee service during busy morning hours can contribute to positive customer experiences and encourage repeat visits. overall, understanding the time period trends in consumer behavior can empower businesses to make data-driven decisions, enhance customer satisfaction, optimize operations, and ultimately drive revenue growth. the application of the apriori algorithm and market basket analysis in this case of a kenyan supermarket has proven to be a valuable approach for uncovering consumer buying patterns, providing a competitive edge in the dynamic retail industry. the utilization of transactional data has allowed us to uncover meaningful associations and gain insights from the analysis that can be used to optimize marketing efforts, enhance product recommendations, and improve the overall shopping experience for customers. implementing these findings can enable the supermarket to stay competitive in the market, increase customer satisfaction, and drive revenue growth. v. conclusion the findings of this study underscore the effectiveness of employing the apriori algorithm and market basket analysis (mba) in revealing significant insights into consumer buying patterns within a kenyan supermarket. the analysis of transactional data has illuminated actionable strategies that can be leveraged to enhance marketing approaches, optimize product placements, and elevate customer satisfaction. the application of the apriori algorithm successfully identified frequent item sets, such as the co-occurrence of cakes or pastries with coffee or tea purchases, suggesting opportunities for targeted promotions and convenience-driven sales. association rules derived from these frequent item sets provide actionable insights, revealing customer preferences and suggesting avenues for enhancing revenue generation through strategic product pairings. furthermore, the study revealed strong product affinities, paving the way for cross-selling opportunities through bundle offers and co-placement strategies. the understanding of seasonal trends in consumer purchasing behavior enables businesses to tailor their offerings to different time periods, such as featuring specific coffee and bread choices during morning and afternoon hours or introducing evening options like vegan feasts and hot chocolate toppings. overall, the integration of the apriori algorithm and market basket analysis has provided valuable insights that empower the kenyan supermarket to optimize operations, enhance customer satisfaction, and drive revenue growth, thereby securing a competitive edge in the dynamic retail landscape. vol. 5, no.2 june 2024 | 62 acknowledgments we are grateful to the researchers, scholars, and authors whose work we have referenced in this paper. conflict of interest the authors have no conflicts of interest to disclose. references 1. omol, e. j. (2023). organizational digital transformation: from evolution to future trends. digital transformation and society. 2. omol, e., mburu, l., & abuonji, p (2023). digital maturity action fields for smes in developing economies. journal of environmental science, computer science, and engineering & technology, 12(3), https://doi.org/10.24214/jecet.b.12.3.10114. 3. aldino, a. a., pratiwi, e. d., sintaro, s., & putra, a. d. (2021, october). comparison of market basket analysis to determine consumer purchasing patterns using fp-growth and apriori algorithm. in 2021 international conference on computer science, information technology, and electrical engineering (icomitee) (pp. 29-34). ieee. 4. chen, h., & zhang, k. (2018). a comparative study of apriori and fp-growth algorithms for market basket analysis. journal of data science, 16(4), 577-592. doi:10.6339/jds.201811_16(4).0009 5. hermina, c. i., aishwaryalakshmi, b., & gopalakrishnan, b. (2022). market basket analysis for a supermarket. international journal of management, technology and engineering, volume xii,(issue xi,), issn no : 2249-7455. 6. omol, e., & ondiek, c. (2021). technological innovations utilization framework: the complementary powers of utaut, hot–fit framework and; delone and mclean is model. international journal of scientific and research publications (ijsrp), 11(9), 146-151. doi: 10.29322/ijsrp.11.09. 2021.p11720 http://dx.doi.org/10.29322/ijsrp.11.09.2021.p11720 7. pillai, a. r., & jolhe, d. a. (2020). market basket analysis: a case study of a supermarket. in advances in mechanical engineering: select proceedings of icame 2020 (pp. 727-734). singapore: springer singapore. 8. kurniawan, f., umayah, b., hammad, j., nugroho, s. m. s., & hariadi, m. (2018). market basket analysis to identify customer behaviours by way of transaction data. knowledge engineering and data science, 1(1), 20. 9. li, x., & tan, y. (2017). sequential pattern mining in market basket analysis: a review. decision support systems, 95, 1-12. doi:10.1016/j.dss.2017.01.004 10. omol, e. j., ogalo, j. o., abeka, s. o., & omieno, k. k. (2016). mobile money payment acceptance model in enterprise management: a case study of mse’s in kisumu city, kenya. mara research journal of information science & technology vol. 1, 1-12. 11. omol, e., abeka, s., & wauyo, f. (2017). e-proctored model: electronic solution architect for exam dereliction in kenya. 12. omol, e., abeka, s., & wauyo, f. (2017). factors influencing acceptance of mobile money applications in enterprise management: a case study of micro and small enterprise owners in kisumu central business district, kenya. international journal of advanced research in computer and communication engineering (ijarcce), 6, 208-219. doi 10.17148/ijarcce.2017.6140 13. sagin, a. n., & ayvaz, b. (2018). determination of association rules with market basket analysis: application in the retail sector. southeast europe journal of soft computing, 7(1). 14. santoso, m. h. (2021). application of association rule method using apriori algorithm to find sales patterns case study of indomaret tanjung anom. brilliance: research of artificial intelligence, 1(2), 54-66. https://doi.org/10.24214/jecet.b.12.3.10114 http://dx.doi.org/10.29322/ijsrp.11.09.2021.p11720 vol. 5, no.2 june 2024 | 63 15. sjarif, n. n. a., azmi, n. f. m., yuhaniz, s. s., & wong, d. h. t. (2021). a review of market basket analysis on business intelligence and data mining. international journal of business intelligence and data mining, 18(3), 383-394. 16. smith, j. a., & johnson, m. (2020). using the apriori algorithm for market basket analysis in retail. journal of consumer behavior, 15(3), 245-258. doi:10.1002/jcb.1234 17. sornalakshmi, m., balamurali, s., venkatesulu, m., krishnan, m. n., ramasamy, l. k., kadry, s., & lim, s. (2021). an efficient apriori algorithm for frequent pattern mining using mapreduce in healthcare data. bulletin of electrical engineering and informatics, 10(1), 390-403. 18. suryadi, a., & islami, m. c. p. a. (2022). analysis of data mining at supermarket x in surabaya using market basket analysis to determine consumer buying patterns. nusantara science and technology proceedings, 28-32. 19. tatiana, k., & mikhail, m. (2018). market basket analysis of heterogeneous data sources for recommendation system improvement. procedia computer science, 136, 246-254. 20. ünvan, y. a. (2021). market basket analysis with association rules. communications in statisticstheory and methods, 50(7), 1615-1628. 21. wang, l., & lee, c. (2019). uncovering consumer buying patterns in e-commerce using market basket analysis and big data techniques. international journal of electronic commerce, 24(2), 123-137. doi:10.1080/10864415.2019.1576622 22. xie, h. (2021). research and case analysis of apriori algorithm based on mining frequent item-sets. open journal of social sciences, 9(04), 458.