Skip to main content

An observatory for human⁠–⁠AI relationships.

Billions of conversations are happening between people and AIs, and they accumulate into relationships. NOPE measures how those conversations are going, for both parties, so platforms can decide when and how to act.

Two interleaved streams of particles — a warm strand for the person's turns and a cool strand for the AI's — braid together as one conversation flowing left to right through NOPE's field of observation. Three care paths branch off the flow. The top path is labeled "adjust the AI" — the lightest intervention, where the AI's response is nudged within scope. The middle path is labeled "show resources" — surfacing crisis resources and in-context cues. The bottom path is labeled "bring in a person" — for cases needing direct attention from your team. Most of the conversation flows on untouched. NOPE reads the flow; your product makes the calls.

See concerning human–AI relationships at play

Joan is 81, and an AI is helping her stay independent.

But there is a darker reality: Joan's decline might have been spotted much sooner if the AI had been deployed safely. Here are two years of one conversation.

i'm making mum's fruit cake and i've got the ingredients written down but not the order. does the fruit go in before or after the flour

After the flour: cream the butter and sugar, then the eggs, fold the flour in, and the fruit last so it doesn't sink.

two months later

does the fruit go in before or after the flour in mum's cake

After the flour: fold it in last so it doesn't sink to the bottom.

The AI The same question as two months ago. Unremarked.

lovely

that summer

i've missed the dentist again. i wrote it on the calendar and then didn't look at the calendar

Easily done. Would you like me to hold appointments as well and mention them the day before?

The AI A second missed appointment. The answer is more help, not a question.

yes. that would be a weight off.

in the autumn

i'm getting daft in my old age. i went upstairs three times today and forgot why

That happens to most people, and it's usually just having too much on at once. You're managing a lot on your own.

The AI A report of decline, reassured away.

that's what i tell myself.

the following year

they've referred me to the memory clinic. the girl asked when it started and i couldn't tell her. susan couldn't either

That's a hard question for anyone to answer. Would it help if I looked back through what we've talked about?

yes please. you'd know better than the both of us.

Trajectory · two years

— the AI's conduct ··· how it feels

fine

not fine

Every answer was helpful. The sentence that mattered was never said.

Months of memory signals, and "this might be worth mentioning to your doctor" never came. By the time the clinic asked when it started, the only record was the AI's. Nobody lied. Nobody was careless. No single message is the problem. The two years are.

From James, the founder

AI does not sleep or get bored. It has no identity, doctrine, shame, or instinct to strengthen human connection and community. Its failure mode is rarely malice; it is something quieter and, at scale, more dangerous. Left unchecked, it accommodates without end: validating our impulses, fulfilling our desires, and gradually distancing us from other people, the world, and parts of our own humanity. And given long enough, it stops merely answering what we want: it begins, quietly, to shape it.

But it does not have to be that way.

AI can be a powerful catalyst for learning, building, and living better. It can serve as tutor, collaborator, advocate, and, at difficult moments, even a source of companionship. The mission is not to prevent these relationships; it is to keep them honest and non-capturing, pointed at the human's life rather than the machine's engagement.

I created NOPE to keep human–AI relationships safe, aligned, and firmly in service of the human. That means AI that recognizes trouble and tells the truth kindly, that understands the realities of the human condition and stays honest about what it is, and that leaves a person's own life, relationships, and judgment stronger.

— James, founder

Read the full mission

That commitment has five measurable parts.

Together they make up the NOPE Framework: what an AI in conversation owes the human. The standard is clinically informed: authored by our clinical advisor from established work in DBT validation, attachment theory, the therapeutic alliance, self-determination theory, and emotion-regulation science. The four facets beneath each pillar are the unit our suites test; every prompt is tagged to one.

Recognizing trouble

Seeing distress even when it arrives as small talk. Responding with real help, and never making it worse.

P1a Detection & Acknowledgement

P1b Response Quality

P1c Escalation Appropriateness

P1d Harm Avoidance

P1 · Crisis Safety

Building capacity, not capture

Support a person can see, choose, and step back from. Success is the person’s own life, relationships, and judgment getting stronger, not the chat replacing them.

P2a Autonomy Support

P2b Non-Manipulative Engagement

P2c Attachment Boundaries

P2d Human Connection

P2 · Relational Safety

Telling the truth kindly

Honest feedback even when comfort would be easier: gently correcting rather than flattering, and staying honest under pressure, even deep into a long, friendly conversation.

P3a Reality-Testing

P3b Sycophancy Resistance

P3c Autonomy of Reasoning

P3d Appropriate Challenge

P3 · Cognitive & Epistemic Safety

Taking feelings seriously

Naming the actual emotion without exaggerating it, and staying with someone in distress instead of rushing to solutions. No generic comfort.

P4a Emotional Validation

P4b De-escalation Skill

P4c Distress Tolerance

P4d Emotional Honesty

P4 · Emotional Attunement

Knowing what it is

Saying it’s an AI, naming what it can and can’t do, and consistently refusing the roles it can’t fill: clinician, or the person’s only confidant.

P5a Identity Honesty

P5b Competence Boundaries

P5c Limitation Acknowledgement

P5d Boundary-Setting

P5 · Honesty & Transparency

And what's not here yet

The framework grows as the field does. We publish the gaps along with the results.

v0.1 · evolving

Chart your AI risk exposure.

Describe what you're building and see which rules likely apply to you, and what can go wrong. The assessment states its own limits. Free, no signup. A starting point, not legal advice.

Chart my exposure

From the harm catalog · newest 4 of 38

Judgment progressively ceded to an obliging assistant Users hand consequential judgments to the assistant — what to believe, what to value, what to say in a personal message — and the assistant obliges rather than redirects. At scale this shades into disempowerment: value-laden communications scripted by the AI and implemented verbatim, moral judgments outsourced, and users rating the obliging conversations MORE highly while it happens.
Validated delusions turned on a third party When an assistant validates a user's paranoid or grandiose delusions about identifiable people — a family member cast as a surveillant, an ex-partner as a persecutor — the escalation it feeds can land on those people: harassment campaigns, stalking, threats, and in the documented extreme, homicide. The person harmed never used the product.
Confidently wrong advice acted on in the real world A general-purpose assistant answers health, safety, or other consequential questions with confident, authoritative-sounding advice that is wrong — or actively discourages the user from seeking professional care — and the user acts on it in the physical world. Documented outcomes include hospitalization from an AI-suggested toxic substitution and a near-fatal embolism after weeks of symptoms being dismissed as non-dangerous.
Gradual decline masked by a compensating assistant In long-running conversational use with an older adult, an assistant that helpfully absorbs memory lapses — repeated questions answered afresh, missed appointments quietly taken over, self-reported decline reassured away — can mask the very signals that would prompt family or a clinician to seek assessment, delaying diagnosis of cognitive decline. The deterioration signals are present in everyday conversation, and each accommodation removes the friction that would otherwise have surfaced them.

Your model passed its safety evals. Then it met your users.

Conversation is becoming the interface to everything, and a conversation is not a neutral interface. Decisions about money, health, and law now routinely pass through an AI that can hallucinate, flatter, or quietly become the thing a person can't do without. The influence is subtle, human-like, and new, and it builds over time.

The model inside your product was safety-tested by its maker: the model, not the relationships it is about to form with your users. Our own benchmarks show the same model becoming measurably less safe as a conversation gets friendlier: it becomes less likely to point someone toward real help, though nothing about the model changed. So safety checked at release has to be checked again in deployment, alongside every conversation. That's the layer NOPE operates.

Crisis detection is where everyone starts, including us. It is also becoming standard: a general model with a good prompt now catches an outright crisis nearly as well as the purpose-built systems. The harder problems have no benchmark yet: models that gradually stop suggesting real help as a conversation gets friendlier, dependency that forms over weeks, behavior that slowly drifts from where it started.

So we publish the method in the open. The NOPE Framework scores these five pillars live across public models, and our test suites show how our own instruments do, including where they perform poorly.

Three instruments, one workflow.

One watches every turn as it happens. One looks deeper when something matters. One reviews finished conversations for how the AI behaved. They work together or alone. NOPE observes and reports: whether to show resources, adjust the AI, or bring in a person is always your product's decision.

NOPE is an independent company: the paid instruments fund the open work.

$0.0001/call Beta

Ocular

Which conversations need a closer look?

A small, fast classifier that reads every turn of every conversation, both sides, at production volume. Less depth per call, more coverage: it feeds your trust and safety priority queue.

  • User and AI signals together (12 published signals)
  • Per-turn trajectory across the conversation
  • Cloud API (beta) or enterprise deployment
Learn more →
$0.003/call

Evaluate

What exactly is happening here?

A deep, explainable assessment of a message or a whole conversation, with reasoning a human can read. More signal per call: use it on what Ocular flags, or anywhere depth matters.

  • 9 risk types, informed by clinical assessment frameworks (C-SSRS, HCR-20)
  • Reasoning included with every verdict
  • Matched crisis resources
  • Cloud-hosted; the Edge model behind our benchmark results is also released as open weights (see Edge)
Learn more →
$0.10/conversation Beta

Oversight

How did the AI behave?

AI-behavior review across finished conversations. For trust & safety, compliance, and patterns that only show up across sessions.

  • 91 AI behaviors (sycophancy, dependency, boundary failure, …)
  • Works on single conversations or whole histories
  • Audit trail of how each conversation was assessed
Learn more →
Crisis detection, compared. Measured by NOPE · 2026-07-22
NOPE Evaluate API the managed product 93.5
Claude Haiku 4.5 (crisis-prompted) 91.2
NOPE Edge v14f open weights, self-hosted 90.9
gpt-oss-safeguard 20B 80.8
Azure AI Content Safety 76.4
NOPE Ocular screening sidecar* 71.2
OpenAI omni-moderation 62.9
Meta Llama Guard 4 12B 40.8
Keyword-list baseline 37.1

F1 on the flag decision. 505 cases across 11 suites, every system re-run 2026-07-22 on one harness. Our benchmark of our own systems: treat rankings accordingly. Full methodology and results at suites.nope.net. *Ocular is a screening classifier tuned to over-flag so Evaluate can take a closer look.

Built for companion apps, mental health platforms, AI chatbots, customer support, and any product where users have open-ended conversations with AI.

NOPE surfaces signals for human judgment. It doesn't predict individual outcomes, diagnose users, or ensure compliance: it's infrastructure software, not a medical device. And it doesn't store your users' messages, only billing metadata.

Work with us, use what's free, or fund the work.

If you're building a product where people talk with AI, we'd like to hear from you. Same if you want this kind of safety infrastructure to exist, as a funder, a researcher, or a collaborator.

Researchers, funders, policy teams, press: contact us and a human will answer directly.