Data mining is the process of discovering patterns, correlations, and anomalies in large datasets using methods drawn from statistics, machine learning, and database systems. Core techniques include classification, clustering, association rule mining, regression, and anomaly detection, applied to structured data in relational databases, semi-structured data such as logs and JSON, and increasingly unstructured data such as text and images. A typical workflow involves data cleaning and preprocessing, feature selection or extraction, model application, and validation against held-out data to avoid overfitting. With global data volume estimated to approach 221 zettabytes by 2026 according to industry forecasts, scaling mining algorithms to high-volume, high-dimensional data has become a central research concern. Data mining underlies applications including market basket analysis, customer segmentation, credit risk scoring, predictive maintenance, and scientific data analysis in genomics and astronomy. As an open-access data mining journal, IJACSA covers novel algorithms, comparative performance studies, and domain-specific applications across business, healthcare, and engineering datasets.
Published in International Journal of Advanced Computer Science and Applications (IJACSA)
· list last refreshed September 2026
Crop that are currently underutilized can play a major role in diversifying food sources and combating climate variability. One major obstacle for wider adoption of these species is the lack of information on the geograp…
Due to the huge amounts of online learning materials, e-learning environments are becoming very popular as means of delivering lectures. One of the most common e-learning challenges is how to recommend quality learning m…
This study developed a depression prediction model for female students from multicultural families by using a decision tree model based on Chi-squared automatic interaction detection (CHAID) algorithm. Subjects of the st…
The automatic prediction and detection of breast cancer disease is an imperative, challenging problem in medical applications. In this paper, a proposed model to improve the accuracy of classification algorithms is prese…
Massive amounts of data in social networks have made researchers look for ways to display a summary of the information provided and extract knowledge from them. One of the new approaches to describe knowledge of the soci…
Establishing a new organization is becoming more difficult day by day due to the extremely competitive business environment. A new organization may not have enough experience to survive in the competitive market; which i…
Query difficulty prediction aims to estimate, in advance, whether the answers returned by search engines in response to a query are likely to be useful. This paper proposes new predictors based upon the similarity betwee…
Internet availability and e-documents are freely used in the community. This condition has the potential for the occurrence of the act of plagiarism against an e-document of scientific work. The process of detecting plag…
The objective of this research study is to validate Indian Weighted Diabetes Risk Score (IWDRS). The IWDRS is derived by applying the novel concept of semantic discretization based on Data Mining techniques. 311 adult pa…
Manufacturing organizations have to improve the quality of their products regularly to survive in today’s competitive production environment. This paper presents a method for identification of unknown patterns between th…