LLMCompass: Evaluation and Resource-Aware Model Selection
Jan 1, 2026
·
1 min read
University of Central Florida — ongoing
Research question: how can an AI system select a model that satisfies task-specific capability requirements while jointly accounting for quality, uncertainty, latency, memory, throughput, and monetary or computational cost?
- Curate skill-oriented evaluations for summarization, question answering, reasoning, instruction following, and related capabilities; compare rankings across datasets, prompts, judges, and deployment settings.
- Study generalization of model rankings to unseen tasks and evaluate whether intent-to-skill mappings provide reliable evidence for model recommendation.
- Develop reproducible inference and data infrastructure as a byproduct of the research, including model execution, metadata capture, ranking, and deployment-plan generation.

Authors
Postdoctoral Researcher, BridgeAI Lab
I’m a postdoctoral researcher in the BridgeAI Lab at the University of Central
Florida, where I study how language models and AI systems should be evaluated,
selected, and deployed when aggregate benchmark scores do not capture semantic
behavior, user requirements, or deployment constraints. I completed my Ph.D. in
Computer Science and Software Engineering at Auburn University in 2025, advised by
Dr. Sathyanarayanan N. Aakur and co-advised by
Dr. Shubhra Kanti (Santu) Karmaker, with a dissertation
on evaluating semantic and contextual alignment in language models. My current work
combines model evaluation, resource-aware model selection, and scalable inference
into reproducible methods for matching AI models to real-world tasks.