All frameworks
Prompting & Evaluation

Prompting & Evaluation

Learn to design effective LLM prompts and build reliable evaluations to measure, compare, and improve model performance.

Linux CLILinux CLI
Module 1

Prompt Engineering

Master the craft of prompting: write prompts that hit accuracy targets on test sets — classification, extraction, format control, few-shot, and chain-of-thought.

6

Ticket Classification Prompt

easy0 / 1 solved

Invoice Field Extraction Prompt

easy1 / 1 solved

House Style Enforcement Prompt

easyNo attempts yet

Few-Shot Uplift

easyNo attempts yet

Reasoning Versus Direct

easyNo attempts yet

Triage Desk Prompt

easyNo attempts yet
Module 2

Evaluation

Learn to measure AI quality: build metrics, scorers, and calibrated LLM judges that agree with humans and catch regressions.

6

Exact Match and F1 Metrics

easy0 / 1 solved

Pairwise Comparison Scorer

easy1 / 1 solved

Answer Quality Judge Rubric

easyNo attempts yet

Regression Hunt

easyNo attempts yet

Judge Calibration

easyNo attempts yet

Release Gate

easyNo attempts yet