[ICLR 2025] xFinder: Large Language Models as Automated Evaluators for Reliable Evaluation
-
Updated
Nov 14, 2025 - Python
[ICLR 2025] xFinder: Large Language Models as Automated Evaluators for Reliable Evaluation
CiteME is a benchmark designed to test the abilities of language models in finding papers that are cited in scientific texts.
Latxa: An Open Language Model and Evaluation Suite for Basque
An evaluation suite for Retrieval-Augmented Generation (RAG).
Single-GPU post-training pipeline for Qwen3-8B math reasoning. Combines QLoRA SFT, DPO, and GRPO with verifiable rewards (correctness + format + length). Seven Colab notebooks, +6pp GSM8K from SFT alone
Industrial-grade benchmarking engine for AI agents. Define test scenarios in YAML, run high-performance parallel evaluations, and generate premium glassmorphism reports. Supports LangGraph, CrewAI, AutoGen, and custom agent stacks.
`decon`, but with python API binding.
Benchmarking SLMs on structured extraction, RAG Q&A, intent classification, and latency
Evaluate models and compare their scores
Practical eval: accuracy, perplexity, simple task probes; harness cmds.
LLM Model: Fine-tuning, Evaluation, Containerization, Deployment, CI/CD Pipeline
To associate your repository with the lm-evaluation topic, visit your repo's landing page and select "manage topics."