Skip to content

Evidnc · 2024

EvidncFounder, Designer & Engineer2024

AI-powered survey analysis, 40 hours down to 15 minutes

Turning 40 hours of qualitative research synthesis into 15 minutes, live at evidnc.ai

Metric
research synthesis time
40hrs → 15min
beta testers ran real sessions
10
shipped at evidnc.ai
Live
person design + engineering
1

Founder, Designer & Engineer · Evidnc

Evidnc · Design + Engineering

Evidnc: AI-powered survey analysis, 40 hours down to 15 minutes

Solo build, 2024. Live product.

02 The 30-second version

The 30-second version

The challenge

Qualitative researchers spend 40+ hours per project manually coding transcripts, finding themes, and building synthesis reports. Existing tools offered keyword search, which is useless for semantic meaning. A researcher saying "I feel lost" and another saying "the navigation confused me" are the same insight that share zero keywords. Every existing platform missed this.

What I did

Built a live AI platform that ingests research transcripts, generates embeddings, clusters them semantically, and surfaces patterns researchers would otherwise miss. The interface is designed around confidence and control: every cluster shows the quotes that formed it, the similarity scores, and a drag-to-reorganize interaction so the researcher is always the final decision-maker. Trust comes from visibility, not from hiding the machinery.

My role

  • Sole designer and engineer
  • Owned research, product, UI, and the full technical stack
  • React frontend, embedding pipeline, semantic clustering algorithm
  • Recruited and ran 10 beta testers through real research synthesis tasks

The team

  • Anirudh PalaskarFounder, Designer & Engineer

The key insight

I assumed researchers wanted better organization tools: tags, folders, smart search. But talking to them revealed they wanted pattern detection. They weren't looking for help arranging what they already knew. They wanted the tool to surface things they'd missed. That reframed everything. The design decision and the architecture decision became the same decision: embeddings, not keywords.

The constraints

  • Solo buildOne person owning research, design, frontend, and the embedding + clustering pipeline. Every feature had to justify its build cost.
  • Trust over accuracyResearchers will not hand their data to a black box. The product had to earn trust before being allowed to help.
  • Ship, not prototypeThe bar was a live, usable product, not a Figma flow. Every screen had to work against real transcripts, not lorem ipsum.

03 Key decisions

Key decisions

  1. 01

    Embeddings over keywords

    Themes in qualitative data are semantic, not lexical. "I feel lost" and "the navigation confused me" are the same insight sharing zero keywords. Every existing tool I looked at was solving this with TF-IDF or string matching: fast, cheap, and wrong. I built an embedding-based clustering pipeline so the system understands meaning, not words. This was simultaneously a design decision (dramatically better results for researchers) and an architecture decision (embeddings over TF-IDF). For Evidnc, those two decisions collapsed into one.

    Embeddings over keywords
  2. 02

    Show the AI's reasoning

    Researchers don't trust black boxes, and they're right not to. I designed every cluster to expose its own evidence: the source quotes that formed it, the similarity scores between them, and a drag-to-reorganize affordance so the researcher can override the model at any point. The clustering UI isn't a magic box that spits out themes, it's a lens that shows its work. Trust came from visibility, not from accuracy claims.

    Show the AI's reasoning
  3. 03

    Ship it, don’t prototype it

    The easy version of Evidnc would have been a Figma prototype and a pitch deck. But researchers evaluate tools by running their actual data through them, not by imagining what a tool could do. I committed to shipping a live product (real ingest, real embeddings, real clustering, real export) before showing it to anyone. That bar forced every design decision to hold up against real transcripts, and it’s why 10 beta testers ran real sessions on their own research instead of walking through a canned demo.

    Ship it, don’t prototype it

The deep dive

04 Context

Context

Qualitative researchers spend 40+ hours per project manually coding survey responses and interview transcripts. That means reading every quote, tagging themes, and reassembling insights into a synthesis report. It’s tedious, slow, and the most valuable part of the work (pattern detection) is the part humans are weakest at.

Every existing tool I evaluated solved this with keyword search or lexical tagging. But themes in qualitative data are semantic, not lexical. A respondent saying "I feel lost" and another saying "the navigation confused me" are expressing the same insight and share zero keywords. TF-IDF misses it. String matching misses it. Even "smart search" misses it, because it was never designed to detect meaning.

I built Evidnc as a live, usable product rather than a prototype, because researchers evaluate tools by running their real data through them, not by imagining what a tool could do.

The problemQualitative research synthesis is manual, slow, and blind to semantic similarity. Researchers need pattern detection, not better organization.

Where it hurt

  • 40+ hours per synthesisManual coding of transcripts, theme-building, and report assembly dominates the researcher’s week.
  • Keyword search misses meaningExisting tools rely on lexical match. Semantically identical quotes go uncounted because they share no words.
  • Black-box AI is worseOff-the-shelf LLM summarizers produce confident answers with no traceable evidence. Researchers can’t defend findings to stakeholders.
  • Context switching kills flowResearchers jump between transcript tool, spreadsheet, Miro board, and doc. Evidence lives in four places, synthesis lives nowhere.

05 Research

Research

Discovery calls with the ASU UX research team

  1. Synthesis is not a single tool problem. It’s a context-switching problem across 4 to 5 disconnected tools.
  2. The most time-consuming step is open-text coding, and it gets worse as researchers fatigue through long transcripts.
  3. Researchers already distrust AI tools they’ve tried. Not because of accuracy, but because the tools couldn’t explain their own outputs when stakeholders asked "why did you categorize this that way?"
  4. The real problem isn’t speed. It’s the gap between what the AI decided and what the researcher can defend.

I used an AI tool to pre-code themes. When a stakeholder asked why a specific response was categorized the way it was, I had no answer. I had to fall back on manual review anyway.

UX Researcher, ASU discovery interview

Open-ended conversations with UX researchers on the ASU teamObserved how they synthesize survey and interview data todayCaptured the tool stack they patch together (Qualtrics, Atlas.ti, Miro, NVivo, Dovetail, Notion)

Quantitative validation survey with 300 researchers

  1. 89.7% of respondents demand full traceability to source data. The strongest signal in the study (Q9).
  2. 92.7% say surfacing ambiguous or low-confidence cases would increase their trust in the tool (Q12).
  3. 76.9% validated the "AI transparency and trust gap" as the #1 composite pain point, outranking even the open-text coding burden.
  4. 77% said they’d pay for a transparent auto-coding tool. Uniform across UXRs, PMs, and VPs, with no segment differences.
  5. 72.2% feel a qual-quant disconnect. They can’t easily connect what respondents said to who said it. Universal, not segment-specific.
  6. Large effect size on AI trust by segment (Cohen’s d = 0.91). PMs adopt AI faster (M=3.75) while UXRs (2.49) and VPs (2.48) need transparency before they’ll adopt at all.

I’d automate first-pass categorization and confidence scoring. Give me a sorted list: high-confidence auto-codes that I approve in bulk, and low-confidence edge cases that I review one by one.

UX Researcher, validation survey

15-question instrument: 12 Likert items plus 3 open-ended promptsUX Researchers, Product Managers, VP / CX LeadersAnalyzed with ANOVA and Cohen’s d to surface segment-level differences and effect sizesUsed to rank pain-point severity and validate hypotheses before locking the product direction

Beta testing with 10 researchers on live product

  1. Researchers didn’t ask "how accurate is the clustering?" They asked "why did you group these together?" That validated the transparency-over-accuracy design call.
  2. The drag-to-reorganize interaction was used more than any AI-suggested action. Control matters more than automation.
  3. Running real transcripts surfaced edge cases no prototype would have: messy quotes, off-topic responses, language mixing. The product held up because it was live, not mocked.
  4. The 40-hour-to-15-minute claim held across all 10 sessions on real datasets.

Recruited 10 researchers and gave them access to the live Evidnc productAsked them to run real research synthesis tasks on their own transcriptsObserved where they got stuck, what they over-trusted, and what they ignoredIterated the product between sessions based on recurring friction points

06 Design pillars

Design pillars

  • 01Semantic over LexicalEmbeddings, not keywords. The system should understand that two sentences mean the same thing even when they share no words.
  • 02Transparency as TrustEvery AI output must expose its evidence. No black boxes. Clusters show source quotes, similarity scores, and let the researcher override them.
  • 03Researcher Stays in ControlThe AI proposes; the researcher disposes. Drag-to-reorganize, manual override, and human-in-the-loop editing at every step.
  • 04Ship, Don’t PrototypeReal ingest, real embeddings, real export. The product had to hold up against real transcripts, not a curated demo dataset.

07 User flows

User flows

Thematic AnalysisI want the tool to find the themes I’d miss reading line-by-line.

  1. 1Upload survey responses or interview transcripts
  2. 2System generates semantic embeddings per response
  3. 3Clusters form automatically based on meaning
  4. 4Review, rename, or merge clusters with drag-to-reorganize
Thematic Analysis

Semantic SearchI want to search my data by idea, not by exact words.

  1. 1Type a concept or question in natural language
  2. 2System ranks responses by semantic similarity, not keyword match
  3. 3Results include quotes that share meaning but no vocabulary
  4. 4Click any result to jump to its source context
Semantic Search

Question Analysis with VisualsI want to see how responses to one question break down.

  1. 1Pick any question in the survey
  2. 2See themed breakdown of all responses to that question
  3. 3Visual charts show distribution of sentiment and themes
  4. 4Drill into any segment to read source quotes
Question Analysis with Visuals

Quick AI SummaryI need a one-paragraph summary I can drop into a report right now.

  1. 1Trigger summary generation on any cluster, question, or full dataset
  2. 2System generates a concise narrative with inline quote citations
  3. 3Every claim is backed by a linked source quote
  4. 4Copy-paste straight into Notion, Docs, or a slide
Quick AI Summary

09 Reflections

Reflections

  1. The design decision was the architecture decision

    Choosing embeddings over TF-IDF wasn’t just a technical optimization. It was the difference between a tool that works and a tool that doesn’t. On Evidnc, the product’s value and the backend’s architecture collapsed into a single call. Owning both roles made that call obvious. I suspect it would have been a months-long debate if design and engineering were separate.

  2. Transparency beats accuracy

    Researchers didn’t ask "how accurate is the clustering?" They asked "why did you group these together?" I stopped chasing a higher silhouette score and instead made every cluster explain itself. Trust came from visibility, not from metrics I could have put on a landing page.

  3. Ship so researchers can test on their own data

    A prototype would have let me show ten beta testers a pretty demo. Shipping a live product let them upload their own transcripts, and that’s when the feedback got useful. Every interesting thing I learned came from someone running their real work through it.

10 Next in the file

Next projectHeyPoco2025HeyPocoVoice-first life logger with a full RAG pipelineOpen the case study