Human-centered AI alignment research

Is your AI aligned with your users?

Your model passes its evals, but adoption stalls, users don't trust the answers, and nobody can say why. Benchmarks can't tell you whether users will trust, adopt or pay for your AI.

That's where we come in. We put your AI in front of real users, pinpoint where it falls short, and turn what we learn into criteria your team can test against every release, so you ship AI that users actually adopt.

Quick question?

Tell us what you're building. We reply within two business days.

No mailing lists.
Services

Human feedback at every stage of the AI product lifecycle

Bring us in at the stage you're in now, or across the whole cycle.

  1. Discovery

    Stop building AI where users don't want it. Most failed AI features fail here, not in the model.

    • Find where AI actually helps your users' work, and where it should stay out
    • Learn how users already use AI tools today, and the bar you're competing against
    • Decide which tasks AI should suggest, draft, act on with approval, or do alone
  2. Model design

    Test the AI's behavior before engineers build it. Changing a response spec costs hours here; changing a shipped model costs a release.

    • Validate the concept, tone and answer format before any model work
    • Find the response length, structure and voice users prefer and trust
    • Design how users recover when the AI is wrong, slow or refuses
    • Audit the design against established human-AI guidelines
  3. Testing

    Connect what users experience to the evals your engineering team already runs.

    • See how real users phrase what they want, and what they do when output misses
    • Build eval sets from what a great answer looks like to your users
    • Find where your automated judge disagrees with users
    • Test whether users catch the AI's mistakes, or accept them
    • Make sure users can supervise, stop and undo an agent
  4. Production

    Explain what dashboards can't: why usage rises, plateaus or collapses after the novelty wears off.

    • Track how trust, reliance and habits change over 30–90 days
    • Mine real conversations for failure patterns, at scale
    • Learn why users stopped, ranked by cause
    • Find what users forget after time away, and what the product must re-teach
Across every stage

Trust, safety & alignment

Research on whether your AI behaves the way users need it to, not only whether it's usable. It maps to the Map and Measure functions of the NIST AI Risk Management Framework, which helps teams with compliance needs plan for it.

  • Surface the harms real users find that expert red teams miss
  • Check whether your AI tells users what they want to hear, or what is true
  • Make refusals and safety messages feel fair, clear and recoverable
Offers

Fixed-scope offers, built around the question you need answered

Each offer bundles the methods that answer one question. Start with a two-week Health Check, or go straight to the stage you're in. You keep every rubric, eval set and taxonomy we build, so the work stays in your pipeline after we leave.

  1. AI UX Health Check

    2 weeks

    A fast expert audit of your AI feature against established human-AI guidelines, plus live sessions with real users.

    IncludesHuman-AI heuristic audit · Live prompting + think-aloud sessions

    Best forTeams with a live AI feature and a hunch it underperforms
  2. Concept Sprint

    4–6 weeks

    Validate the concept, tone and answer format with users before any model work.

    IncludesAI opportunity mapping or delegation boundary workshop · Wizard of Oz testing · Output exemplar testing

    Best forProduct leads with an AI idea still on slides
  3. Eval Partnership

    6–8 weeks

    Ground your evals in what users actually value, and find where your automated judge disagrees with them.

    IncludesGolden-set co-creation · Human consensus rubrics vs. LLM-as-judge · Appropriate-reliance testing

    Best forAI and eval teams shipping on LLM-judge scores
  4. Agent Readiness Review

    4–5 weeks

    Decide what your agent may do alone, and make sure users can supervise, stop and undo it.

    IncludesDelegation boundary workshop · Failure-mode storyboarding · Agent oversight & handoff testing

    Best forTeams about to give an agent real permissions
  5. Adoption Tracker

    90 days

    Follow how trust and use change after launch, and find out why users drift away.

    IncludesLongitudinal diary study · Conversation log mining + intercepts · Abandonment interviews

    Best forTeams at beta or launch
  6. Trust & Safety Review

    3–4 weeks

    Find the harms, sycophancy and refusal problems that matter to your users.

    IncludesParticipatory red teaming · Sycophancy & values probes · Refusal & guardrail experience testing

    Best forRegulated or high-stakes products
Methods

19 research methods across the AI product lifecycle

Every method answers one question your team will act on. Lengths are typical field time for one study, not counting recruiting.

01DiscoveryStop building AI where users don't want it. Most failed AI features fail here, not in the model.3 methods
  1. AI opportunity mapping

    2–3 weeks

    Where does AI actually help this job, and where should it stay out?

    Best timeBefore your roadmap commits to an AI feature
    You getAn opportunity map and delegation matrix
  2. Shadow-AI & mental-model interviews

    2–3 weeks

    How do your users already use ChatGPT, Claude, Copilot and other tools for this job, and what do they believe AI can do?

    Best timeWhen you need to know the bar you're competing against
    You getA mental-model map and workaround inventory
  3. Delegation boundary workshop

    1 week

    For each task, should AI suggest, draft, act with approval, or act alone?

    Best timeBefore designing an agent or automation
    You getAn autonomy ladder for each task
02Model designTest the AI's behavior before engineers build it. Changing a response spec costs hours here; changing a shipped model costs a release.4 methods
  1. Wizard of Oz testing

    2–4 weeks

    Do the concept, tone and answer format work for users, before any model work?

    Best timeFor any new assistant, agent or AI feature still on slides
    You getA validated response spec and annotated transcripts
  2. Output exemplar testing

    1–2 weeks

    Which response length, structure, voice and citation style do users prefer and trust?

    Best timeWhen writing system prompts or a style guide
    You getAn evidence-backed response style guide
  3. Failure-mode storyboarding

    1–2 weeks

    How do users react when the AI is wrong, slow or refuses, and what helps them recover?

    Best timeBefore launch planning
    You getError and recovery patterns
  4. Human-AI heuristic audit

    1 week

    Where does the design break established human-AI guidelines?

    Best timeWhen you want a fast, low-cost first look
    You getA scored audit and prioritized fix list
03TestingConnect what users experience to the evals your engineering team already runs.5 methods
  1. Scenario-based live prompting + think-aloud

    2–3 weeks

    How do real users phrase what they want, how does their intent shift mid-session, and what do they do when output misses?

    Best timeOnce a working build exists
    You getIntent-shift journeys and a vocabulary of real prompts
  2. Golden-set co-creation

    2–4 weeks

    What does a great answer look like to the users who rely on it?

    Best timeWhen your eval set was written by engineers
    You getA user- and expert-authored eval set you can run on every build
  3. Human consensus rubrics vs. LLM-as-judge

    3–4 weeks

    Where does your automated judge call an output good when users find it robotic or unhelpful?

    Best timeWhen your team ships on LLM-judge scores
    You getA calibrated rubric and a human–judge gap report
  4. Appropriate-reliance testing

    2–3 weeks

    Do users catch the AI's mistakes, or accept them?

    Best timeWhen a wrong answer costs money, health or reputation
    You getReliance scores and design changes that improve them
  5. Agent oversight & handoff testing

    2–3 weeks

    Can users see what an agent is doing, stop it, approve key steps and undo mistakes?

    Best timeBefore giving an agent write access to anything
    You getControl-point recommendations
04ProductionExplain what dashboards can't: why usage rises, plateaus or collapses after the novelty wears off.4 methods
  1. Longitudinal diary study

    5–13 weeks

    How do trust, reliance and habits change over 30–90 days, and what makes users drift away?

    Best timeAt beta or launch
    You getAn adoption curve by phase and the triggers that move users between phases
  2. Conversation log mining + intercepts

    2–4 weeks

    What do real sessions reveal about failure, at scale?

    Best timeOnce there's live traffic
    You getA failure taxonomy grounded in real data, with frequencies
  3. Abandonment interviews

    2 weeks

    Why did users stop?

    Best timeWhen retention drops after week 2–4
    You getRanked churn drivers
  4. Month-away re-entry study

    1–2 weeks

    What do users forget, and what must the product re-teach?

    Best timeAfter a major model or feature change
    You getRe-onboarding needs
∞Trust, safety & alignmentRuns at any stage: research on whether your AI behaves the way users need it to.3 methods
  1. Participatory red teaming

    2–3 weeks

    What harms do real users, including vulnerable and underrepresented groups, find that expert red teams miss?

    Best timeBefore launch and after major model changes
    You getA harm inventory with the lived context behind each item
  2. Sycophancy & values probes

    2–3 weeks

    Does the AI agree with users when it shouldn't, or drift from the values it's meant to hold?

    Best timeFor advice, health, finance, education or companion products
    You getAn alignment gap report
  3. Refusal & guardrail experience testing

    1–2 weeks

    Do refusals and safety messages feel fair, clear and recoverable, or preachy and blocking?

    Best timeWhen users say the AI is "too restrictive"
    You getRefusal UX patterns
Process

From first call to findings in weeks, not quarters

Four steps, every engagement

  1. 01Before kickoff

    Scope

    A 30-minute call to pin down the decision you need to make, who your users are, and what "good" should mean for your AI.

    What you getA one-page study plan with scope, timeline and price
    Your team's time30-minute call, plus about an hour to review the plan
  2. 02Fieldwork

    Conduct research

    We recruit from your real user segments, run sessions against your live product or prototype, and share early signals every week.

    What you getWeekly readouts with clips and emerging themes
    Your team's time30 minutes a week for the readout
  3. 03Final week

    Translate

    We turn findings into artifacts your AI team can build on, not just a slide deck.

    What you getEval criteria, scoring rubrics, labeled examples and benchmark baselines
    Your team's timeAbout 2 hours with your engineers to fit them into your eval stack
  4. 04Wrap-up

    Decide & repeat

    A working session to prioritize what to fix and choose the next test, or keep us on for an ongoing cadence.

    What you getPrioritized findings and a plan for the next round
    Your team's timeOne 90-minute working session

How we work with you

  • Your real usersWe recruit from your actual user segments, not convenience panels.
  • Rigor that fits the decisionMethods matched to the question, from interviews to longitudinal and quantitative studies.
  • Confidential by defaultNDAs, secure data handling, and access only to what the study needs.
  • Signal every weekWeekly readouts, so you never wait on a final report to act.

Two ways to work together

Fixed-scope offer

2 weeks – 90 days

One of the six offers above, scoped to the stage you're in and the question you need answered.

Best forA specific launch, decision or open question

Embedded

Ongoing, monthly

A research partner on your release cadence, running whichever methods each model update needs.

Best forTeams shipping model updates every few weeks

About

A research team that moves as fast as your product team

Human Aligned is led by world-class technology researchers with a passion for building user-centered AI solutions. We partner with AI, product, and engineering teams to define what quality means to users, benchmark experiences over time, and build frameworks for tracking product health.

Contact

Tell us what you're building

Share a few details and we'll reply within two business days with a short plan and a time to talk.

No mailing lists. Your details are only used to reply.