Fixed-scope offer
2 weeks – 90 daysOne of the six offers above, scoped to the stage you're in and the question you need answered.
Best forA specific launch, decision or open question
Your model passes its evals, but adoption stalls, users don't trust the answers, and nobody can say why. Benchmarks can't tell you whether users will trust, adopt or pay for your AI.
That's where we come in. We put your AI in front of real users, pinpoint where it falls short, and turn what we learn into criteria your team can test against every release, so you ship AI that users actually adopt.
Bring us in at the stage you're in now, or across the whole cycle.
Stop building AI where users don't want it. Most failed AI features fail here, not in the model.
Test the AI's behavior before engineers build it. Changing a response spec costs hours here; changing a shipped model costs a release.
Connect what users experience to the evals your engineering team already runs.
Explain what dashboards can't: why usage rises, plateaus or collapses after the novelty wears off.
Research on whether your AI behaves the way users need it to, not only whether it's usable. It maps to the Map and Measure functions of the NIST AI Risk Management Framework, which helps teams with compliance needs plan for it.
Each offer bundles the methods that answer one question. Start with a two-week Health Check, or go straight to the stage you're in. You keep every rubric, eval set and taxonomy we build, so the work stays in your pipeline after we leave.
A fast expert audit of your AI feature against established human-AI guidelines, plus live sessions with real users.
IncludesHuman-AI heuristic audit · Live prompting + think-aloud sessions
Validate the concept, tone and answer format with users before any model work.
IncludesAI opportunity mapping or delegation boundary workshop · Wizard of Oz testing · Output exemplar testing
Ground your evals in what users actually value, and find where your automated judge disagrees with them.
IncludesGolden-set co-creation · Human consensus rubrics vs. LLM-as-judge · Appropriate-reliance testing
Decide what your agent may do alone, and make sure users can supervise, stop and undo it.
IncludesDelegation boundary workshop · Failure-mode storyboarding · Agent oversight & handoff testing
Follow how trust and use change after launch, and find out why users drift away.
IncludesLongitudinal diary study · Conversation log mining + intercepts · Abandonment interviews
Find the harms, sycophancy and refusal problems that matter to your users.
IncludesParticipatory red teaming · Sycophancy & values probes · Refusal & guardrail experience testing
Every method answers one question your team will act on. Lengths are typical field time for one study, not counting recruiting.
Where does AI actually help this job, and where should it stay out?
How do your users already use ChatGPT, Claude, Copilot and other tools for this job, and what do they believe AI can do?
For each task, should AI suggest, draft, act with approval, or act alone?
Do the concept, tone and answer format work for users, before any model work?
Which response length, structure, voice and citation style do users prefer and trust?
How do users react when the AI is wrong, slow or refuses, and what helps them recover?
Where does the design break established human-AI guidelines?
How do real users phrase what they want, how does their intent shift mid-session, and what do they do when output misses?
What does a great answer look like to the users who rely on it?
Where does your automated judge call an output good when users find it robotic or unhelpful?
Do users catch the AI's mistakes, or accept them?
Can users see what an agent is doing, stop it, approve key steps and undo mistakes?
How do trust, reliance and habits change over 30–90 days, and what makes users drift away?
What do real sessions reveal about failure, at scale?
Why did users stop?
What do users forget, and what must the product re-teach?
What harms do real users, including vulnerable and underrepresented groups, find that expert red teams miss?
Does the AI agree with users when it shouldn't, or drift from the values it's meant to hold?
Do refusals and safety messages feel fair, clear and recoverable, or preachy and blocking?
A 30-minute call to pin down the decision you need to make, who your users are, and what "good" should mean for your AI.
We recruit from your real user segments, run sessions against your live product or prototype, and share early signals every week.
We turn findings into artifacts your AI team can build on, not just a slide deck.
A working session to prioritize what to fix and choose the next test, or keep us on for an ongoing cadence.
One of the six offers above, scoped to the stage you're in and the question you need answered.
Best forA specific launch, decision or open question
A research partner on your release cadence, running whichever methods each model update needs.
Best forTeams shipping model updates every few weeks
Human Aligned is led by world-class technology researchers with a passion for building user-centered AI solutions. We partner with AI, product, and engineering teams to define what quality means to users, benchmark experiences over time, and build frameworks for tracking product health.
Share a few details and we'll reply within two business days with a short plan and a time to talk.