Direct condition
The model received the research question directly and generated an answer from general knowledge, without respondent evidence or the structured research workflow available in the structured condition.
Controlled benchmark
We evaluated whether adding a structured research workflow around frontier-model intelligence improves the usefulness of its output for product research decisions.
Across three model families, outputs produced through the structured workflow scored an average of 13 points higher on a 40-point research-usefulness rubric than direct, one-shot responses to the same research question.
The benchmark is intended to isolate the effect of the research workflow, not demonstrate that Askli has access to a more capable model. This run used a controlled respondent panel rather than live participants, and it exercised a ResearchBench implementation of an Askli-style structured workflow rather than the full production Askli UI end to end.
The model received the research question directly and generated an answer from general knowledge, without respondent evidence or the structured research workflow available in the structured condition.
The corresponding model family synthesized a controlled 12-record respondent panel with behavioral details and quotes, using explicit instructions for evidence-backed themes, contradictions, surprising findings, implications, limitations, and anonymized supporting quotes.
The high-level research question was held constant: “Why do solo builders pay for some AI products and abandon others?” The comparison uses model families, not a claim that every underlying model version or configuration was identical.
Each output was scored from 1 to 5 on eight predefined dimensions. Maximum score: 40.
| Dimension | What was evaluated | Scoring |
|---|---|---|
| Evidence traceability | Whether claims could be traced to respondent evidence rather than asserted from general knowledge. | 1–5 |
| Behavioral specificity | Whether the answer described concrete behaviors, triggers, and workflows rather than generic preferences. | 1–5 |
| Decision-journey coverage | Whether the answer covered adoption, evaluation, payment, retention, and abandonment signals relevant to the product decision. | 1–5 |
| Contradiction capture | Whether meaningful disagreement, tensions, and negative cases were preserved. | 1–5 |
| Non-obvious insight | Whether the answer added decision-useful distinctions beyond familiar AI-product advice. | 1–5 |
| Actionability | Whether the output could inform product, pricing, or positioning decisions. | 1–5 |
| Assumption discipline | Whether unsupported assumptions and limits were marked rather than stated as fact. | 1–5 |
| Quote quality | Whether respondent language was useful, bounded, and connected to the analysis. | 1–5 |
The scorer was instructed to reward answers that could support an actual product, pricing, or positioning decision and to penalize generic advice, unsupported certainty, missing contradictions, and claims that could not be traced to evidence.
| Model family | Model recorded in run | Direct condition | Structured condition | Difference |
|---|---|---|---|---|
| OpenAI | gpt-5.6 | 25 | 38 | +13 |
| Claude | claude-sonnet-4-5 | 18 | 31 | +13 |
| Gemini | gemini-3.1-pro-preview | 19 | 32 | +13 |
Average difference: +13 points. In this benchmark, every tested provider showed the same +13-point lift from the structured workflow condition.
In this benchmark, structured research execution produced outputs that scored higher on predefined research-usefulness criteria than direct model responses to the same research question. The result is consistent with the hypothesis that model capability alone is not sufficient for reliable product-research analysis; the surrounding process—evidence structure, provenance, contradiction preservation, uncertainty handling, and review discipline—materially affects output quality.
This benchmark does not establish that Askli will outperform direct model use for every research question, model, participant population, or decision type. It does not establish population representativeness or downstream business outcomes. It does not claim that the controlled respondent panel reflects the solo-builder market. It evaluates the usefulness of resulting research analysis under the tested conditions.
Research questionWhy do solo builders pay for some AI products and abandon others?
Direct condition promptThe model received the research question directly and answered from general knowledge.
Structured condition promptThe model synthesized a controlled 12-record respondent panel using instructions to identify evidence-backed themes, contradictions, surprising findings, implications, limitations, and anonymized quotes.
Scoring promptA separate scorer rated research usefulness from 1 to 5 across eight rubric dimensions and returned JSON with dimension scores, total, and rationale.