Validate the featurebefore you build it.
ShipSift sends a team of research agents across podcasts, job postings, case studies and pricing pages, checks every quote against its source, and has competing AI models react as your buyer personas. You get a verdict, a scorecard and the questions to ask real customers, in about 5 to 7 minutes.
- 01Research
- 02Evidence
- 03Triage
- 04Jury
- 05Red team
- 06Judge
- 07Cross-check
Findings
RESHAPE · 0/100
Sample run: a seven-pass replay ending in the verdict RESHAPE, 46 out of 100, with 78% evidence. 104 calls, 0 failed, 5 minutes 31 seconds.
Models and sources ShipSift orchestrates
One question. Models and sources from:
Trademarks belong to their owners. ShipSift is not affiliated with or endorsed by them.
The problem
The people who could validate your roadmap won't take your call
Buyers, users and industry insiders rarely pick up for a founder. Getting access takes a network, a research budget, or weeks of chasing. Most teams guess, ask one chatbot, or ship and find out.
Founders
You need proof before you build. A week of chasing calls is a week of runway.
Product owners
You need to know which feature matters to which buyer, and which ones nobody would switch for.
Product marketing and GTM
You need personas that are evidenced, not invented in a doc nobody trusts.
ShipSift reads what your buyers already say in public, then argues about it so you don't have to guess.
Use cases
Three ways to validate, before you build or sell
Pick the question. ShipSift runs the same evidence pipeline behind each one.
Is there demand, who pays for what today, and what does it cost? You get a verdict: go, reshape or park.
RESHAPE · 46/100
Recommendation
Reshape before you build. The evidence supports part of the idea and undercuts the rest, with the counter-case attached.
Every claim behind the score carries its quote and its source.
Sample run · real output layout
How it works
One question in, one decision out
Ask about an idea, a feature or a persona. ShipSift fans out across deep web research, targeted search, full-page reads and podcast transcripts, then converges on one memo. Drag the playhead to replay a real run.
The run, in order: Research agent, Evidence agents, Classifier agent, Jury, Red team, Judge, Cross-check agent. Research took 1 minute 37 seconds, evidence 1 minute 1 second, triage 3 seconds and the jury 1 minute 32 seconds, for a total of 5 minutes 31 seconds.
The agent swarm
One question. 104 calls. 0 failed.
Look at everything one question sets off. Every call is a node; every pulse below is one of them, fired in pipeline order.
Calls by role in run 1 (1 persona): 1 Research agent, 1 Search agent, 7 Page-read & podcast agents, 18 Model calls, 77 Classifier calls. In run 2 (3 personas): 1 Research agent, 3 Search agent, 9 Page-read & podcast agents, 26 Model calls, 78 Classifier calls. Run 1 read 35 pieces of material in 12 batches, and 30 podcast segments; 44 podcast claims were extracted and 41 kept, 3 dropped because the quote was not found; 57 evidence claims gave 54 kept, 3 dropped (5%); 15 of 15 jury scores were re-checked.
The 7-pass workflow
Seven passes, each logged with its own health report
Not a chatbot: separate stages for research, evidence, triage, a jury, a red team, a judge and an independent cross-check.
- Research (Research agent): A deep web research agent finds what's actually been published: prices, customers, case studies, job postings. Figures missing from snippets get re-checked on the full page.
- Evidence (Evidence agents): Search, page reads and podcast episodes, with a quote pulled for every claim. A claim whose quote isn't in the source is dropped.
- Triage (Classifier agent): A classifier labels each finding by relevance, persona, job and stance, and flags any figure its source doesn't contain.
- Jury (Jury): Three different AI models react in character as each of your buyer personas. Their reactions are simulated, and labeled that way.
- Red team (Red team): A separate model builds the strongest case against the idea.
- Judge (Judge): A frontier model writes the memo: a verdict per persona, a scorecard, and the questions to ask real people.
- Cross-check (Cross-check agent): An independent classifier re-scores the jury. Disagreements are shown, not hidden.
Multi-model jury
No single model gets the last word.
Three models from three different labs play your buyers. A fourth argues against you. A fifth writes the verdict. Classifier agents check the work, and if a model is down, a named fallback steps in.
Jury: Mistral Large 3, GPT-6 Astra and Grok 4.7. Red team: Kimi K3. Judge: Claude Opus 5.5. Classifier agents handle triage and cross-check, with named fallbacks.
Where the models disagreed
Simulated reactions · one questionA gap finder “would save hours by consolidating competitor pre-launch drops.”
“pre-book is about what competitors book six months out, not what is live.”
“Useful for investigation,” but it “needs demonstrably better coverage”
Live run log
Every run leaves a receipt
A replay of a real run: what ran, what it kept, what it dropped and why.
$ shipsift run --persona "Merchandiser · Mid-market" --question "Would $5M+ Shopify apparel brands pay for sourced competitor intel?" research agent: 3 pages re-checked ✓ 1m 37s evidence agents: 5 searches · 1 page read ✓ podcast agent: 3 topics → 30 podcast segments podcast claims: 44 extracted · 41 kept · 3 dropped (quote not found) extraction: 57 claims · 3 dropped (5%) · 54 kept ✓ 1m 1s classifier agent: relevance · persona · job · stance · 1 figure flagged ✓ jury: 3 models × persona 3/3 ✓ red team: counter-case written ✓ judge: memo · verdict: RESHAPE · number check clean ✓ cross-check agent: 15/15 re-scored ✓ run complete · 104 calls · 0 failed · 5m 31s
Auto-classification
Every finding, sorted before you read it.
Relevance, persona, job-to-be-done, stance and source, labeled automatically. The part that used to eat your evenings. Anything a source doesn't back is flagged, not smoothed over.
- RELEVANT 0.88: Labeled by relevance, persona, job and stance. (via Research agent)
- Single source (Podcast): Brands still benchmark competitor prices by hand, over 3–4 weeks. (via Podcast agent · single source)
- UNDERCUTS 0.89: Published plans cap tracked competitors at 5–10. (via Search agent)
- Vendor claim: Labeled as marketing, weighted accordingly. (via Search agent)
- UNSUPPORTED FIGURE: 152000000: flagged, not in any source behind this finding. (via Classifier agent)
Measured on a real run
You get the output, not just a number
A weighted scorecard, the verdict it adds up to, and what the cross-check disagreed on.
Weighted score
46/100
RESHAPE · evidence 78%
- Job coverage55 (25%)
- Workflow fit50 (20%)
- Data trust45 (20%)
- Switching cost40 (15%)
- Willingness to pay35 (20%)
- Cross-check disagreed on 2 of 15 items. Shown, not hidden.
- Number check: every figure in the memo appears in a source.
- passes per run
- 7
- per run
- 5–7 min
- calls in two runs, 0 failed
- 104 and 117
- claims dropped: quote not found
- 3–5%
- cross-check re-scored
- 100%
Old way vs ShipSift
Validate this week, not next quarter
Drag the handle. Same question, two ways of answering it.
The old way
- Chase prospects for discovery calls that don't get booked
- Pay for a market report that's already dated
- Ask one chatbot and hope
- Guess at personas in a doc nobody trusts
- Ship the feature and find out
With ShipSift
- Ask tonight, read a cited answer in minutes
- Buyer voices from podcasts, job postings and case studies
- Competing models play your personas, labeled as simulated
- A jobs × features map before you write a spec
- Walk into real calls with the exact questions to ask
Trust
Built so you can check its work
Every number checked
Each figure in the memo must appear in a source behind the finding it cites, or it's flagged in plain sight.
No silent failures
If a step comes back short, the run says so and offers a one-click retry of only what's missing.
Evidence and simulation, kept apart
Cited findings are real; jury reactions are clearly labeled as simulated.
Fallbacks, named
If a model refuses or is down, a backup steps in and the run page says which.
FAQ
Questions people ask first
Tell us what you're trying to decide.
ShipSift is working with a small number of teams. Tell us the question and we'll take it from there.