All case studies

Feedback Analytics Platform · Fine-tuned LLM classification

Fine-tuning a feedback taxonomy classifier to 93.9% F1

The problem

Thousands of daily feedback items needed tagging against a 5-level taxonomy. Aspect classification across unrelated topics collapsed to near-random performance at F1 0.19.

What changed

  • Theme–subtheme classifier live at 93.9% F1, 99.8% valid JSON
  • Aspect classifier improved 3.5× (0.19 → 0.67) via calibration alone
  • Validated path to 0.85+ F1 with no new data collection