LLMCompass: Evaluation and Resource-Aware Model Selection

Jan 1, 2026 · 1 min read
projects

University of Central Florida — ongoing

Research question: how can an AI system select a model that satisfies task-specific capability requirements while jointly accounting for quality, uncertainty, latency, memory, throughput, and monetary or computational cost?

  • Curate skill-oriented evaluations for summarization, question answering, reasoning, instruction following, and related capabilities; compare rankings across datasets, prompts, judges, and deployment settings.
  • Study generalization of model rankings to unseen tasks and evaluate whether intent-to-skill mappings provide reliable evidence for model recommendation.
  • Develop reproducible inference and data infrastructure as a byproduct of the research, including model execution, metadata capture, ranking, and deployment-plan generation.
Yash Mahajan
Authors
Postdoctoral Researcher, BridgeAI Lab
I’m a postdoctoral researcher in the BridgeAI Lab at the University of Central Florida, where I study how language models and AI systems should be evaluated, selected, and deployed when aggregate benchmark scores do not capture semantic behavior, user requirements, or deployment constraints. I completed my Ph.D. in Computer Science and Software Engineering at Auburn University in 2025, advised by Dr. Sathyanarayanan N. Aakur and co-advised by Dr. Shubhra Kanti (Santu) Karmaker, with a dissertation on evaluating semantic and contextual alignment in language models. My current work combines model evaluation, resource-aware model selection, and scalable inference into reproducible methods for matching AI models to real-world tasks.