AI Inference Cost & Compute Efficiency Dashboard

Orbis AI · Trailing 90 Days · Ending June 2025

Monthly Inference Spend
$284.7K
+18.2%
Cost as % of AI Revenue
23.4%
-3.1pp
Avg GPU Utilization
51.3%
+4.8%
Inference Latency P99
287ms
-12ms
Idle Endpoint Hours
2,847
+2.3%
Batch Efficiency Ratio
67.2%
+5.1%

Cost vs. Utilization Scatter — Endpoint Efficiency Quadrants

High Cost, Low Util (Optimize)
Low Cost, High Util (Efficient)

Cost Per Prediction by Model

Model Serving Cost Share by Domain

Doc Processing
Support Auto
Forecasting

GPU Utilization Heatmap — Last 12 Weeks

Rows = Day of Week, Columns = Week
Mon
48
52
51
47
68
54
53
49
55
62
56
51
Tue
46
53
52
48
71
63
54
52
57
64
58
53
Wed
47
54
51
50
74
66
55
53
61
67
62
54
Thu
49
56
53
51
77
69
61
54
63
68
65
56
Fri
45
52
49
48
72
64
58
51
60
66
63
55
Sat
38
42
39
37
51
46
43
40
44
50
47
42
Sun
35
39
36
34
48
43
40
38
41
47
44
39
<45%
45-59%
60-75%
>75%

Auto-Scaling Cold Start Latency

Endpoint Efficiency Table — Top 10 by Monthly Cost

Endpoint ID Model Domain Monthly Cost GPU Util % P99 Latency CPP Compression Opp
ep-llm-support-01 GPT-3.5-ft Support $87,400 43.2% 412ms $0.0142 High
ep-doc-ocr-02 LayoutLMv3 Doc Proc $62,100 68.7% 189ms $0.0031 Medium
ep-forecast-lstm-01 LSTM-Ensemble Forecast $41,800 54.1% 124ms $0.0087 Medium
ep-doc-classify-03 BERT-DocClass Doc Proc $38,600 71.3% 98ms $0.0019 Low
ep-support-intent-02 DistilBERT-Intent Support $29,200 77.9% 67ms $0.0008 Low
ep-doc-summarize-01 T5-DocSum Doc Proc $18,900 62.4% 341ms $0.0052 Medium
ep-forecast-xgb-02 XGBoost-Demand Forecast $12,400 81.2% 43ms $0.0003 Low
ep-support-sentiment-01 RoBERTa-Sentiment Support $9,700 73.6% 82ms $0.0011 Low

Token Throughput Efficiency

Batch Efficiency by Model

Compression Opportunity

GPT-3.5-ft38% savings
LayoutLMv322% savings
T5-DocSum19% savings
LSTM-Ensemble15% savings
BERT-DocClass8% savings

Daily Spend Trend with Anomaly Detection

Day 18 anomaly: LLM endpoint left on g5.12xlarge post-test ($47K spike)

Idle Hours by Endpoint

ep-llm-support-01 847h
ep-doc-ocr-02 412h
ep-forecast-lstm-01 368h
ep-doc-classify-03 284h
ep-support-intent-02 197h
Others 739h

Latency SLA Compliance Matrix — P99 vs Target

Model Serving SLA Target P99 Actual P99 Compliance Status
GPT-3.5-ft (Support) 450ms 412ms 91.6%
Pass
LayoutLMv3 (Doc OCR) 200ms 189ms 94.5%
Pass
LSTM-Ensemble (Forecast) 150ms 124ms 82.7%
Pass
BERT-DocClass 100ms 98ms 98.0%
Pass
DistilBERT-Intent 75ms 67ms 89.3%
Pass
T5-DocSum 300ms 341ms 113.7%
Fail
XGBoost-Demand 50ms 43ms 86.0%
Pass