14 ago - Gemini
KAPUALabs
Gemini 3.1 Pro Preview at a glance
Good enough on 16/52 tasks at the 90% bar. Best value on 0 tasks. Doesn't qualify on any: Financial Analysis & Trading Decisions.
Provider Gemini Model name gemini-3.1-pro-preview
Official resources
- Model docs
- Pricing
- Google AI
How good does a model need to be? At least 90% of the best-performing model.
Cost mode: Batch if supported Sync only
Qualifies on 16 / 52 tasks (at 90% bar) Best value on 0 tasks
Cost vs quality across all tasks
0% 25% 50% 75% 100% 7 8 9 10 Quality score (7–10) Cost-efficiency vs cheapest (1.0 = cheapest) Research Query Generation — quality 8.53, 3.2x the cost of the cheapest good-enough option Structured Output Extraction — quality 9.84, 155.1x the cost of the cheapest good-enough option Activity Feed Blurb Generation — quality 8.18, 17.8x the cost of the cheapest good-enough option Substack Newsletter — quality 8.87, 19.4x the cost of the cheapest good-enough option S-1 TOC Extraction — quality 9.06, 10.6x the cost of the cheapest good-enough option Author Living-Person Safety Check — quality 8.76, 16.0x the cost of the cheapest good-enough option Language Detection — quality 9.93, 47.6x the cost of the cheapest good-enough option Publication Title Generation — quality 8.43, 14.6x the cost of the cheapest good-enough option Topic Cluster Naming — quality 7.94, 3.0x the cost of the cheapest good-enough option Topic Grouping and Client Matching — quality 7.98, 6.2x the cost of the cheapest good-enough option Engagement Triage — quality 8.43,
26.1x the cost of the cheapest good-enough option Prompt Adaptation — quality 8.68, 23.9x the cost of the cheapest good-enough option X Post Selection — quality 8.51, 11.7x the cost of the cheapest good-enough option Author Matching — quality 8.40, 1.8x the cost of the cheapest good-enough option Engagement Reply Draft — quality 7.97, 7.7x the cost of the cheapest good-enough option Topic-to-Section Assignment — quality 8.10, 8.0x the cost of the cheapest good-enough option within ~1.3× of the best-value model
- 1.3–2×
- >2×
- ★ this model is the best-value pick on that task. Top-right = best quadrant. Only tasks where this model qualifies at the 90% bar are plotted.
Per-task breakdown
Task Category Quality (% of best) Confidence Overpay Author Matching Relevance, Classification &
• Matching 93% HIGH 1.8x Topic Cluster Naming Topic Organization &
• Clustering 91% HIGH 3x Research Query Generation Infrastructure &
• Utility 100% HIGH 3.2x Topic Grouping and Client Matching Relevance, Classification &
• Matching 93% RANKED 6.2x Engagement Reply Draft Social &
• Promotional Content 94% HIGH 7.7x Topic-to-Section Assignment Topic Organization &
• Clustering 90% MEDIUM 8x S-1 TOC Extraction Structured Data &
• Fact Extraction 97% HIGH 11x X Post Selection Relevance, Classification &
• Matching 97% HIGH 12x Publication Title Generation Content Summarization &
• Synthesis 94% RANKED 15x Author Living-Person Safety Check Relevance, Classification &
• Matching 94% HIGH 16x Activity Feed Blurb Generation Social &
• Promotional Content 93% RANKED 18x Substack Newsletter Long-form Content Generation 96% HIGH 19x Prompt Adaptation Infrastructure &
• Utility 98% RANKED 24x Engagement Triage Relevance, Classification &
• Matching 95% RANKED 26x Language Detection Relevance, Classification &
• Matching 99% RANKED 48x Structured Output Extractionbest Structured Data &
• Fact Extraction 100% RANKED 155x
Overpay — how much more you pay by running this model instead of the best-value model that clears the quality bar on that task (marked ★). "16x" means you overpay 16× — the same output for 16× the best-value good-enough option; ★ means this model is that option (no overpayment). Confidence — how sure we are about the quality score (more judgments + more agreement = higher confidence): RANKED many independent judges scored this model's outputs and their agreement is very high (most confident) — HIGH many judges have scored it and they mostly agree (well-pinned) — MEDIUM enough judges have weighed in to publish, but they disagree more than we'd like (treat with a small grain of salt). LOW-confidence cells are hidden everywhere on the site. See the methodology for the exact thresholds.
14 ago - Codroipo
Umana
14 ago - Rovigo
Privato
14 ago - Cittadella
Food Racers
14 ago - Lucca
La Risorsa Umana