Podnami

Best performance on Podnami

LLM Rankings

Models from Settings → AI Selection, scored for live co-hosts: speed (TTFT), reliability, conversation quality (CQ), and credit value. High speed alone does not mean high CQ.

Last completed run: 2026-08-18T01:14:17.800955+00:00 · 21 models

Click a column header to sort.

Ranked models

# Model Stars Overall Speed Reliability CQ ? CQ — Conversation Quality
Suitability for Podnami live co-hosts: tool calling, no speaker-label leaks, handoffs, short memory, spoken brevity, an in-character chat check, plus whether the model is a non-thinking chat-oriented model vs a utility/doc or heavy-reasoning model.
Value TTFT p50 Tier Badges
1 ▲3
Mistral Nemo
Mistral · mistralai/mistral-nemo
★★★★★ 97.8 95 100 99 97 0.43s Budget Best value
2
Gemini 2.5 Lite
Google · google/gemini-2.5-flash-lite
🔒 2 runs at #2 ▲ best #1
★★★★★ 97.8 95 100 99 97 0.42s Budget
3 ▲13
DeepseekV32
DeepSeek · deepseek/deepseek-v3.2
★★★★★ 96.2 95 100 95 96 1.37s Budget
4 ▼1
Incllusion Ling 2.6
Inclusion · inclusionai/ling-2.6-1t
▲ best #2
★★★★★ 96.2 95 100 95 96 0.76s Budget
5 ▲10
Amazon Nova Lite V1
Amazon · amazon/nova-lite-v1
▲ best #1
★★★★★ 94.9 100 100 89 96 0.40s Budget
6 ▲4
Meta Llamma Scout
Meta Llamma Scout · meta-llama/llama-4-scout
▲ best #2
★★★★★ 94.5 95 100 91 95 0.27s Budget
7 ▲7
GPT-4o Mini
OpenAI · openai/gpt-4o-mini
▲ best #3
★★★★★ 94.2 95 100 99 73 0.78s Standard
8 ▲5
Claude 3 Haiku
Anthropic · anthropic/claude-3-haiku
▲ best #4
★★★★★ 94.2 95 100 99 73 0.35s Standard
9 ▼3
Claude Haiku 4.5
Anthropic · anthropic/claude-haiku-4.5
▲ best #5
★★★★★ 94.2 95 100 99 73 1.23s Standard
10 ▲2
Google Gemini Flash Lite 3.5
Google · google/gemini-3.5-flash-lite
▲ best #3
★★★★★ 94.0 100 100 95 74 0.46s Standard
11 ▼10
DeepSeek Chat
DeepSeek · deepseek/deepseek-chat
▲ best #1
★★★★★ 93.1 80 100 99 91 1.26s Budget
12 ▼5
GPT-3.5 Turbo
OpenAI · openai/gpt-3.5-turbo
▲ best #4
★★★★★ 92.5 95 100 95 72 0.52s Standard
13 ▼8
Ministral 14b-2512
Ministral 14b-2512 · mistralai/ministral-14b-2512
▲ best #2
★★★★★ 91.4 80 100 95 89 0.45s Budget
14 ▼3
GPT-4 Turbo
OpenAI · openai/gpt-4-turbo
▲ best #11
★★★★★ 91.3 95 100 99 54 1.35s Premium
15 ▼6
Gemma 4 31B IT
Google · google/gemma-4-31b-it
▲ best #1
★★★★★ 89.7 80 100 91 88 0.20s Budget
16 ▼8
Inclusion Ling 3.0 Flash
Inclusion · inclusionai/ling-3.0-flash
▲ best #8
★★★★★ 85.5 65 95 95 82 0.92s Budget
17
DeepSeek V4 Flash
DeepSeek · deepseek/deepseek-v4-flash
🔒 2 runs at #17 ▲ best #12
★★★★☆ 84.5 60 95 96 80 2.44s Budget
18 ▲1
Step 3.5 Flash
StepFun · stepfun/step-3.5-flash
▲ best #17
★★★☆☆ 68.3 60 72 72 67 2.74s Budget
19 ▼1
Inclusion Ling 2.6 Flash
Ling · inclusionai/ling-2.6-flash
▲ best #2
★★★☆☆ 68.3 95 38 65 73 0.56s Budget Unstable

Free options (comparison)

# Model Stars Overall Speed Reliability CQ ? CQ — Conversation Quality
Suitability for Podnami live co-hosts: tool calling, no speaker-label leaks, handoffs, short memory, spoken brevity, an in-character chat check, plus whether the model is a non-thinking chat-oriented model vs a utility/doc or heavy-reasoning model.
Value TTFT p50 Tier Badges
1
FREE LLM
OpenRouter · free-llm-cycler
🔒 2 runs at #1
★★★★☆ 78.6 65 100 77 76 1.20s Free Free option Not recommended for live co-hosts
2
Free LLM Router
Open Router · openrouter/free
🔒 2 runs at #2 ▲ best #1
★★★★☆ 70.5 30 95 86 63 2.71s Free Free option Not recommended for live co-hosts

/ = rank change since the previous run · = no change · NEW = first appearance

Probes are scripted OpenRouter checks (not live mixer sessions). Pass/fail + timings only — no reply transcripts stored. Scores reflect a sampled run (a few attempts per model at a non-zero temperature), so small differences between runs are normal — treat close scores as ties, not regressions.