CERESResearch Repository

​​AI-Driven high-throughput document classification and knowledge distillation using hybrid machine learning and large language models through agentic AI ​

dc.contributor.advisorGuo, Weisi
dc.contributor.author​​Kevogo​, ​​Dennis​
dc.date.accessioned2026-03-12T13:58:06Z
dc.date.available2026-03-12T13:58:06Z
dc.date.freetoread2026-03-12
dc.date.issued2025-09
dc.descriptionTang​, Yun
dc.description.abstractEfficient document classification at enterprise scale is increasingly pertinent as organizations must manage millions of unstructured documents per hour while ensuring compliance, high accuracy, operational efficiency, and transparency. This thesis presents a hybrid AI system that combines traditional machine learning (ML) and agentic AI, leveraging large language models (LLMs), to process and classify over one million documents per hour with high precision. This system also enables knowledge distillation for continual learning and cost control. Another important goal of the system is to repurpose high-value talent away from repetitive classification tasks, allowing them to focus on more strategic and impactful work within the organization. The design blends metadata- and content-based ML with content-aware LLM inference, orchestrated by a smart, self-healing, and adaptive routing agent that assigns classification tasks based on complexity and confidence scores. A comprehensive evaluation against established metrics demonstrates the system’s superiority over monolithic approaches, with empirical results underscoring its robustness, adaptability, and auditable intelligence. Traditional machine learning excels in speed and cost, while recent advances in Large Language Models (LLMs) have made nuanced document understanding feasible. This research presents a hybrid document classification system that integrates machine learning and agentic AI reasoning with a knowledge distillation correction loop, achieving scalable, compliant, and adaptive enterprise document governance. Results demonstrate throughput exceeding 1,000,000 documents/hour, robust classification accuracy, cost-effective LLM orchestration, and continuous improvement driven by user feedback. Comparative benchmarking and multi-metric analysis validate the system’s efficacy, generalizability, and auditability.​
dc.description.coursenameMSc in Applied Artificial Intelligence
dc.identifier.urihttps://dspace.lib.cranfield.ac.uk/handle/1826/25029
dc.language.isoen
dc.publisherCranfield University
dc.publisher.departmentAIRS
dc.subjectDocument Classification
dc.subjectLarge Language Models (LLMs)
dc.subjectHybrid Artificial Intelligence
dc.subjectKnowledge Distillation
dc.subjectEnterprise Data Governance
dc.subjectMachine Learning
dc.subjectAgentic AI
dc.subjectScalability
dc.subjectModel Routing
dc.subjectAdaptive Learning
dc.subjectImplicit Reasoning in LLMs
dc.subjectSelf-Evolving Agents
dc.subjectCost-Aware Routing
dc.subjectReal-time Learning Components
dc.subjectAdaptive routing optimization
dc.title​​AI-Driven high-throughput document classification and knowledge distillation using hybrid machine learning and large language models through agentic AI ​
dc.typeThesis
dc.type.qualificationlevelMasters
dc.type.qualificationnameMSc

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Dennis-Kevogo-2025.pdf
Size:
3.84 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.63 KB
Format:
Item-specific license agreed upon to submission
Description: