AI Model Performance & Inference Quality
NovaMind AI Production Observability
All Models
GPT-4o
Claude 3.5
Llama-3-70B
Last 30 Days
⚠️ Recent Event: Hallucination spike on /generate endpoint (day 14, 5.1%) due to prompt template change. Guardrail deployed, rate recovered to 2.3% by day 18. ECE spike on /classify endpoint (day 18, 0.094) from Llama-3-70B feature drift (PSI 0.27).
Inference Latency P95 Over Time by Endpoint
Calibration Error Curve (Reliability Diagram)
Data Drift Index (PSI) Heatmap
Feature Cluster × Week (Last 4 Weeks)
Week -3
Week -2
Week -1
Current
Doc Type
0.04
0.12
0.27
0.18
Text Length
0.03
0.05
0.06
0.11
Language Mix
0.02
0.03
0.04
0.05
User Segment
0.07
0.14
0.16
0.13
Timestamp
0.01
0.02
0.03
0.04
PSI: ≤0.1 stable, 0.1-0.2 monitor, ≥0.2 retrain
Model Version Traffic Share
Hallucination Rate Trend by Endpoint
Model Performance Comparison Table
| Model Version |
Endpoint |
Traffic % |
ECE |
Latency P95 |
Hallucination % |
Token Efficiency |
Review Escalation % |
Gross Margin/Request |
Status |
| GPT-4o-v1.2 |
/generate |
38.4% |
0.052 |
520ms |
2.3% |
8.9 |
4.2% |
$0.042 |
Healthy |
| GPT-4o-v1.2 |
/classify |
14.2% |
0.048 |
380ms |
1.8% |
9.2 |
3.1% |
$0.051 |
Healthy |
| GPT-4o-v1.2 |
/summarize |
7.4% |
0.059 |
640ms |
3.1% |
7.8 |
5.9% |
$0.038 |
Monitor |
| Claude-3.5-v2 |
/generate |
18.1% |
0.061 |
490ms |
2.7% |
8.4 |
5.3% |
$0.039 |
Monitor |
| Claude-3.5-v2 |
/classify |
6.7% |
0.055 |
410ms |
2.2% |
8.7 |
4.1% |
$0.045 |
Healthy |
| Claude-3.5-v2 |
/summarize |
3.2% |
0.063 |
550ms |
3.4% |
7.6 |
7.2% |
$0.033 |
Monitor |
| Llama-3-70B-v3 |
/generate |
4.8% |
0.071 |
430ms |
3.9% |
7.2 |
9.1% |
$0.062 |
Monitor |
| Llama-3-70B-v3 |
/classify |
5.1% |
0.094 |
320ms |
4.3% |
6.8 |
12.4% |
$0.058 |
Critical |
| Llama-3-70B-v3 |
/summarize |
2.1% |
0.068 |
480ms |
4.1% |
7.1 |
8.7% |
$0.067 |
Monitor |
Token Efficiency Ratio by Model Family
AI Inference Health Timeline (Last 30 Days)