Best performance on Podnami
LLM Rankings
Models from Settings → AI Selection, scored for live co-hosts: speed (TTFT), reliability, conversation quality (CQ), and credit value. High speed alone does not mean high CQ.
Last completed run: 2026-08-18T01:14:17.800955+00:00 · 21 models
Click a column header to sort.
Ranked models
| # | Model | Stars | Overall | Speed | Reliability |
CQ
?
CQ — Conversation Quality Suitability for Podnami live co-hosts: tool calling, no speaker-label leaks, handoffs, short memory, spoken brevity, an in-character chat check, plus whether the model is a non-thinking chat-oriented model vs a utility/doc or heavy-reasoning model. |
Value | TTFT p50 | Tier | Badges |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 ▲3 |
Mistral Nemo
Mistral · mistralai/mistral-nemo
|
★★★★★ | 97.8 | 95 | 100 | 99 | 97 | 0.43s | Budget | Best value |
| 2 – |
Gemini 2.5 Lite
Google · google/gemini-2.5-flash-lite
🔒 2 runs at #2
▲ best #1
|
★★★★★ | 97.8 | 95 | 100 | 99 | 97 | 0.42s | Budget | |
| 3 ▲13 |
DeepseekV32
DeepSeek · deepseek/deepseek-v3.2
|
★★★★★ | 96.2 | 95 | 100 | 95 | 96 | 1.37s | Budget | |
| 4 ▼1 |
Incllusion Ling 2.6
Inclusion · inclusionai/ling-2.6-1t
▲ best #2
|
★★★★★ | 96.2 | 95 | 100 | 95 | 96 | 0.76s | Budget | |
| 5 ▲10 |
Amazon Nova Lite V1
Amazon · amazon/nova-lite-v1
▲ best #1
|
★★★★★ | 94.9 | 100 | 100 | 89 | 96 | 0.40s | Budget | |
| 6 ▲4 |
Meta Llamma Scout
Meta Llamma Scout · meta-llama/llama-4-scout
▲ best #2
|
★★★★★ | 94.5 | 95 | 100 | 91 | 95 | 0.27s | Budget | |
| 7 ▲7 |
GPT-4o Mini
OpenAI · openai/gpt-4o-mini
▲ best #3
|
★★★★★ | 94.2 | 95 | 100 | 99 | 73 | 0.78s | Standard | |
| 8 ▲5 |
Claude 3 Haiku
Anthropic · anthropic/claude-3-haiku
▲ best #4
|
★★★★★ | 94.2 | 95 | 100 | 99 | 73 | 0.35s | Standard | |
| 9 ▼3 |
Claude Haiku 4.5
Anthropic · anthropic/claude-haiku-4.5
▲ best #5
|
★★★★★ | 94.2 | 95 | 100 | 99 | 73 | 1.23s | Standard | |
| 10 ▲2 |
Google Gemini Flash Lite 3.5
Google · google/gemini-3.5-flash-lite
▲ best #3
|
★★★★★ | 94.0 | 100 | 100 | 95 | 74 | 0.46s | Standard | |
| 11 ▼10 |
DeepSeek Chat
DeepSeek · deepseek/deepseek-chat
▲ best #1
|
★★★★★ | 93.1 | 80 | 100 | 99 | 91 | 1.26s | Budget | |
| 12 ▼5 |
GPT-3.5 Turbo
OpenAI · openai/gpt-3.5-turbo
▲ best #4
|
★★★★★ | 92.5 | 95 | 100 | 95 | 72 | 0.52s | Standard | |
| 13 ▼8 |
Ministral 14b-2512
Ministral 14b-2512 · mistralai/ministral-14b-2512
▲ best #2
|
★★★★★ | 91.4 | 80 | 100 | 95 | 89 | 0.45s | Budget | |
| 14 ▼3 |
GPT-4 Turbo
OpenAI · openai/gpt-4-turbo
▲ best #11
|
★★★★★ | 91.3 | 95 | 100 | 99 | 54 | 1.35s | Premium | |
| 15 ▼6 |
Gemma 4 31B IT
Google · google/gemma-4-31b-it
▲ best #1
|
★★★★★ | 89.7 | 80 | 100 | 91 | 88 | 0.20s | Budget | |
| 16 ▼8 |
Inclusion Ling 3.0 Flash
Inclusion · inclusionai/ling-3.0-flash
▲ best #8
|
★★★★★ | 85.5 | 65 | 95 | 95 | 82 | 0.92s | Budget | |
| 17 – |
DeepSeek V4 Flash
DeepSeek · deepseek/deepseek-v4-flash
🔒 2 runs at #17
▲ best #12
|
★★★★☆ | 84.5 | 60 | 95 | 96 | 80 | 2.44s | Budget | |
| 18 ▲1 |
Step 3.5 Flash
StepFun · stepfun/step-3.5-flash
▲ best #17
|
★★★☆☆ | 68.3 | 60 | 72 | 72 | 67 | 2.74s | Budget | |
| 19 ▼1 |
Inclusion Ling 2.6 Flash
Ling · inclusionai/ling-2.6-flash
▲ best #2
|
★★★☆☆ | 68.3 | 95 | 38 | 65 | 73 | 0.56s | Budget | Unstable |
Free options (comparison)
| # | Model | Stars | Overall | Speed | Reliability |
CQ
?
CQ — Conversation Quality Suitability for Podnami live co-hosts: tool calling, no speaker-label leaks, handoffs, short memory, spoken brevity, an in-character chat check, plus whether the model is a non-thinking chat-oriented model vs a utility/doc or heavy-reasoning model. |
Value | TTFT p50 | Tier | Badges |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 – |
FREE LLM
OpenRouter · free-llm-cycler
🔒 2 runs at #1
|
★★★★☆ | 78.6 | 65 | 100 | 77 | 76 | 1.20s | Free | Free option Not recommended for live co-hosts |
| 2 – |
Free LLM Router
Open Router · openrouter/free
🔒 2 runs at #2
▲ best #1
|
★★★★☆ | 70.5 | 30 | 95 | 86 | 63 | 2.71s | Free | Free option Not recommended for live co-hosts |
▲/▼ = rank change since the previous run · – = no change · NEW = first appearance
Probes are scripted OpenRouter checks (not live mixer sessions). Pass/fail + timings only — no reply transcripts stored. Scores reflect a sampled run (a few attempts per model at a non-zero temperature), so small differences between runs are normal — treat close scores as ties, not regressions.
