Controlled benchmark

Benchmark: direct model answers vs. evidence-grounded research workflows

We evaluated whether adding a structured research workflow around frontier-model intelligence improves the usefulness of its output for product research decisions.

Summary

Across three model families, outputs produced through the structured workflow scored an average of 13 points higher on a 40-point research-usefulness rubric than direct, one-shot responses to the same research question.

The benchmark is intended to isolate the effect of the research workflow, not demonstrate that Askli has access to a more capable model. This run used a controlled respondent panel rather than live participants, and it exercised a ResearchBench implementation of an Askli-style structured workflow rather than the full production Askli UI end to end.

What was compared

Direct condition

The model received the research question directly and generated an answer from general knowledge, without respondent evidence or the structured research workflow available in the structured condition.

Structured workflow condition

The corresponding model family synthesized a controlled 12-record respondent panel with behavioral details and quotes, using explicit instructions for evidence-backed themes, contradictions, surprising findings, implications, limitations, and anonymized supporting quotes.

The high-level research question was held constant: “Why do solo builders pay for some AI products and abandon others?” The comparison uses model families, not a claim that every underlying model version or configuration was identical.

Evaluation framework

Each output was scored from 1 to 5 on eight predefined dimensions. Maximum score: 40.

DimensionWhat was evaluatedScoring
Evidence traceabilityWhether claims could be traced to respondent evidence rather than asserted from general knowledge.1–5
Behavioral specificityWhether the answer described concrete behaviors, triggers, and workflows rather than generic preferences.1–5
Decision-journey coverageWhether the answer covered adoption, evaluation, payment, retention, and abandonment signals relevant to the product decision.1–5
Contradiction captureWhether meaningful disagreement, tensions, and negative cases were preserved.1–5
Non-obvious insightWhether the answer added decision-useful distinctions beyond familiar AI-product advice.1–5
ActionabilityWhether the output could inform product, pricing, or positioning decisions.1–5
Assumption disciplineWhether unsupported assumptions and limits were marked rather than stated as fact.1–5
Quote qualityWhether respondent language was useful, bounded, and connected to the analysis.1–5

The scorer was instructed to reward answers that could support an actual product, pricing, or positioning decision and to penalize generic advice, unsupported certainty, missing contradictions, and claims that could not be traced to evidence.

Results

Model familyModel recorded in runDirect conditionStructured conditionDifference
OpenAIgpt-5.62538+13
Claudeclaude-sonnet-4-51831+13
Geminigemini-3.1-pro-preview1932+13

Average difference: +13 points. In this benchmark, every tested provider showed the same +13-point lift from the structured workflow condition.

What this result supports

In this benchmark, structured research execution produced outputs that scored higher on predefined research-usefulness criteria than direct model responses to the same research question. The result is consistent with the hypothesis that model capability alone is not sufficient for reliable product-research analysis; the surrounding process—evidence structure, provenance, contradiction preservation, uncertainty handling, and review discipline—materially affects output quality.

What this result does not establish

This benchmark does not establish that Askli will outperform direct model use for every research question, model, participant population, or decision type. It does not establish population representativeness or downstream business outcomes. It does not claim that the controlled respondent panel reflects the solo-builder market. It evaluates the usefulness of resulting research analysis under the tested conditions.

Benchmark materials and transparency

Research questionWhy do solo builders pay for some AI products and abandon others?

Direct condition promptThe model received the research question directly and answered from general knowledge.

Structured condition promptThe model synthesized a controlled 12-record respondent panel using instructions to identify evidence-backed themes, contradictions, surprising findings, implications, limitations, and anonymized quotes.

Scoring promptA separate scorer rated research usefulness from 1 to 5 across eight rubric dimensions and returned JSON with dimension scores, total, and rationale.

Start with your question