CERESResearch Repository

​​AI-Driven high-throughput document classification and knowledge distillation using hybrid machine learning and large language models through agentic AI ​

Loading...
Thumbnail Image

Date published

Free to read from

2026-03-12

Supervisor/s

Industry supervisor/s

Journal Title

Journal ISSN

Volume Title

Department

AIRS

Type

ISSN

Format

Citation

Abstract

Efficient document classification at enterprise scale is increasingly pertinent as organizations must manage millions of unstructured documents per hour while ensuring compliance, high accuracy, operational efficiency, and transparency. This thesis presents a hybrid AI system that combines traditional machine learning (ML) and agentic AI, leveraging large language models (LLMs), to process and classify over one million documents per hour with high precision. This system also enables knowledge distillation for continual learning and cost control. Another important goal of the system is to repurpose high-value talent away from repetitive classification tasks, allowing them to focus on more strategic and impactful work within the organization. The design blends metadata- and content-based ML with content-aware LLM inference, orchestrated by a smart, self-healing, and adaptive routing agent that assigns classification tasks based on complexity and confidence scores. A comprehensive evaluation against established metrics demonstrates the system’s superiority over monolithic approaches, with empirical results underscoring its robustness, adaptability, and auditable intelligence. Traditional machine learning excels in speed and cost, while recent advances in Large Language Models (LLMs) have made nuanced document understanding feasible. This research presents a hybrid document classification system that integrates machine learning and agentic AI reasoning with a knowledge distillation correction loop, achieving scalable, compliant, and adaptive enterprise document governance. Results demonstrate throughput exceeding 1,000,000 documents/hour, robust classification accuracy, cost-effective LLM orchestration, and continuous improvement driven by user feedback. Comparative benchmarking and multi-metric analysis validate the system’s efficacy, generalizability, and auditability.​

Description

Tang​, Yun

Software description

Software language

Git repository

Keywords

Document Classification, Large Language Models (LLMs), Hybrid Artificial Intelligence, Knowledge Distillation, Enterprise Data Governance, Machine Learning, Agentic AI, Scalability, Model Routing, Adaptive Learning, Implicit Reasoning in LLMs, Self-Evolving Agents, Cost-Aware Routing, Real-time Learning Components, Adaptive routing optimization

DOI

Rights

Funder/s

Grant number

Relationships

Relationships

Resources