LLMCompass: Evaluation and Resource-Aware Model Selection
A research platform for capability-aware and resource-aware selection among more than 50 language models across more than 25 task skills.
I study how language models and AI systems should be evaluated, selected, and deployed when aggregate benchmark scores do not fully capture semantic behavior, user requirements, or deployment constraints. My work combines natural language processing, model evaluation, scalable inference, and agentic systems to develop reliable and reproducible methods for matching AI models to real-world tasks.
Current directions include reliable capability and alignment evaluation, resource-aware model selection and routing, evaluation of multi-step and agentic systems, and efficient multi-GPU inference.
A research platform for capability-aware and resource-aware selection among more than 50 language models across more than 25 task skills.
An automated model-sharding workflow for distributing large language models across multiple GPUs.
Software and data for task-free evaluation of sentence embeddings using five semantic similarity alignment criteria.