Data mining is the process of discovering patterns, correlations, and anomalies in large datasets using methods drawn from statistics, machine learning, and database systems. Core techniques include classification, clustering, association rule mining, regression, and anomaly detection, applied to structured data in relational databases, semi-structured data such as logs and JSON, and increasingly unstructured data such as text and images. A typical workflow involves data cleaning and preprocessing, feature selection or extraction, model application, and validation against held-out data to avoid overfitting. With global data volume estimated to approach 221 zettabytes by 2026 according to industry forecasts, scaling mining algorithms to high-volume, high-dimensional data has become a central research concern. Data mining underlies applications including market basket analysis, customer segmentation, credit risk scoring, predictive maintenance, and scientific data analysis in genomics and astronomy. As an open-access data mining journal, IJACSA covers novel algorithms, comparative performance studies, and domain-specific applications across business, healthcare, and engineering datasets.
Published in International Journal of Advanced Computer Science and Applications (IJACSA)
· list last refreshed September 2026
Stroke is a serious disease that has a significant impact on the quality of life and safety of patients. Accurately predicting stroke risk is of great significance for preventing and treating stroke. In the past few year…
It is of great importance for Higher Education (HE) institutions to continuously work on detecting at-risk students based on their performance during their academic journey with the purpose of supporting their success an…
The main motivation of any educational institution is to provide quality education. Therefore, choosing an academic track can be clearly seen as an obstacle, for students and universities, which in turn led to imposing a…
With the rise in technology, a huge volume of data is being processed using data mining, especially in the healthcare sector. Usually, medical data consist of a lot of personal data, and third parties utilize it for the…
In the context of the 4.0 technology revolution, which develops and applies strongly in many fields, in which the banking sector is considered to be the leading one, the application of algorithms to detect fraud is extre…
The importance of Customer Relationship Management (CRM) has never been higher. Thus, companies are forced to adopt new strategies to focus on customers, given the competitive climate in which they operate. Also, compani…
The practise of recognising unauthorised abnormal actions on computer systems is referred to as intrusion detection. The primary goal of an Intrusion Detection System (IDS) is to identify user behaviours as normal or abn…
Computer network attacks are among the most significant and common threats against computer-wired and wireless communications. Intrusion detection technology is used to secure computer networks by monitoring network traf…
Clustering Spatio-temporal data is challenging because of the complexity of processing the spatial and temporal aspects. Various enhanced clustering approaches, such as partition-based and hierarchical-based algorithms h…
With the emergence of digital newspapers and social media, one can easily suffer from information overload. The enormous amount of data they provide has created several new challenges for computational and data mining, e…