Data mining is the process of discovering patterns, correlations, and anomalies in large datasets using methods drawn from statistics, machine learning, and database systems. Core techniques include classification, clustering, association rule mining, regression, and anomaly detection, applied to structured data in relational databases, semi-structured data such as logs and JSON, and increasingly unstructured data such as text and images. A typical workflow involves data cleaning and preprocessing, feature selection or extraction, model application, and validation against held-out data to avoid overfitting. With global data volume estimated to approach 221 zettabytes by 2026 according to industry forecasts, scaling mining algorithms to high-volume, high-dimensional data has become a central research concern. Data mining underlies applications including market basket analysis, customer segmentation, credit risk scoring, predictive maintenance, and scientific data analysis in genomics and astronomy. As an open-access data mining journal, IJACSA covers novel algorithms, comparative performance studies, and domain-specific applications across business, healthcare, and engineering datasets.
Published in International Journal of Advanced Computer Science and Applications (IJACSA)
· list last refreshed September 2026
Speech is now routine evidence in criminal investigations, but forensic audio rarely matches the clean assumptions of standard speaker recognition. Clips are short, noisy, codec-compressed, and channel-mismatched, and th…
Blockchain technology has emerged as a transformative innovation in digital accounting, offering robust mechanisms to enhance auditability, data integrity, and fraud prevention. This study examines how blockchain-based a…
This study aims to examine the impact of corporate ownership structures on tax avoidance using a predictive data mining approach. The main challenge addressed is understanding how variations in ownership influence a firm…
Frequent itemset mining is a widely adopted data mining technique. The application of the technique can be found in transaction database analysis, such as exploring a set of purchased items. Presently, the growing concer…
Federated Learning (FL) offers a privacy-preserving and decentralized paradigm for machine learning, making it particularly suitable for analyzing sensitive psychological and physiological data. This study aims to develo…
Text mining methods often rely on a single data source or simple word frequency statistics, making it difficult to capture multi-source text semantic associations and local contextual dependencies, resulting in poor mini…
Defective products in manufacturing can be reduced by accurately predicting quality outcomes based on process parameters. This study proposes a quality prediction framework for semiconductor manufacturing using the Data…
Diabetes mellitus presents a growing prevalence at the global level, representing a significant public health challenge. Despite the availability of specific treatments, it is imperative to develop innovative strategies…
One vital key for effective management of cloud resources is the ability to predict their users’ consumption’s patterns in granular level. It can provide more insightful analysis to guide these users towards more resourc…
The rapid development of the Internet of Things (IoT) has highlighted the importance of Wi-Fi sensor networks in efficiently collecting data anytime and anywhere. This paper aims to propose an optimized routing protocol…