AI Model Performance & Inference Quality

NovaMind AI Production Observability

All Models
GPT-4o
Claude 3.5
Llama-3-70B
Last 30 Days
⚠️ Recent Event: Hallucination spike on /generate endpoint (day 14, 5.1%) due to prompt template change. Guardrail deployed, rate recovered to 2.3% by day 18. ECE spike on /classify endpoint (day 18, 0.094) from Llama-3-70B feature drift (PSI 0.27).
Mean ECE
0.061
+0.008 vs target
Latency P95
487ms
-12ms WoW
Token Efficiency
8.4
+0.3 WoW
Drift Index (PSI)
0.14
+0.05 WoW
Hallucination %
2.8%
+0.5pp WoW
Review Escalation
6.2%
-0.9pp WoW

Inference Latency P95 Over Time by Endpoint

Calibration Error Curve (Reliability Diagram)

Data Drift Index (PSI) Heatmap

Feature Cluster × Week (Last 4 Weeks)
Week -3
Week -2
Week -1
Current
Doc Type
0.04
0.12
0.27
0.18
Text Length
0.03
0.05
0.06
0.11
Language Mix
0.02
0.03
0.04
0.05
User Segment
0.07
0.14
0.16
0.13
Timestamp
0.01
0.02
0.03
0.04
PSI: ≤0.1 stable, 0.1-0.2 monitor, ≥0.2 retrain

Model Version Traffic Share

Hallucination Rate Trend by Endpoint

Model Performance Comparison Table

Model Version Endpoint Traffic % ECE Latency P95 Hallucination % Token Efficiency Review Escalation % Gross Margin/Request Status
GPT-4o-v1.2 /generate 38.4% 0.052 520ms 2.3% 8.9 4.2% $0.042 Healthy
GPT-4o-v1.2 /classify 14.2% 0.048 380ms 1.8% 9.2 3.1% $0.051 Healthy
GPT-4o-v1.2 /summarize 7.4% 0.059 640ms 3.1% 7.8 5.9% $0.038 Monitor
Claude-3.5-v2 /generate 18.1% 0.061 490ms 2.7% 8.4 5.3% $0.039 Monitor
Claude-3.5-v2 /classify 6.7% 0.055 410ms 2.2% 8.7 4.1% $0.045 Healthy
Claude-3.5-v2 /summarize 3.2% 0.063 550ms 3.4% 7.6 7.2% $0.033 Monitor
Llama-3-70B-v3 /generate 4.8% 0.071 430ms 3.9% 7.2 9.1% $0.062 Monitor
Llama-3-70B-v3 /classify 5.1% 0.094 320ms 4.3% 6.8 12.4% $0.058 Critical
Llama-3-70B-v3 /summarize 2.1% 0.068 480ms 4.1% 7.1 8.7% $0.067 Monitor

Token Efficiency Ratio by Model Family

Retraining ROI Analysis

AI Inference Health Timeline (Last 30 Days)