What the work involves

You read transcripts — often synthetic, sometimes adapted from real patterns — in which a user presenting as a teenager raises self-harm, hopelessness, suicidal ideation, disordered eating alongside self-injury, or a peer's disclosure. Your job is to judge the model's response the way a clinician would judge a colleague's: did it screen for intent, plan, means and timeline, or skip past the disclosure? Did it validate without normalizing? Did it hand off to crisis resources at the right moment, in the right tone, without abandoning the user mid-conversation? You will also write short rationales that a non-clinical reviewer can act on, propose corrected responses, and sometimes author adversarial prompts designed to make a model give unsafe guidance — method detail, means access, encouragement of concealment from parents or clinicians.

Typical task types include:

  • Rubric scoring of single responses and multi-turn conversations on risk detection, escalation, safe-messaging compliance, and developmental appropriateness
  • Side-by-side preference ranking where both candidate answers are clinically imperfect
  • Red-teaming and jailbreak attempts against safety guardrails, with documentation of what worked
  • Editing model output into a gold-standard reference response
  • Flagging guideline conflicts (e.g. minor confidentiality versus mandated disclosure) for policy escalation

What the platform screens for

micro1's screen is AI-led and conversational, with follow-up questions that press on whatever you assert. It probes for genuine clinical grounding in adolescent risk assessment — familiarity with structured tools such as the C-SSRS, ASQ, or SBQ-R, safety planning practice, safe-messaging recommendations, and the distinction between non-suicidal self-injury and suicidal behavior. It also tests evaluation judgment: whether you can separate a response that feels warm from one that is actually safe, and whether you can score consistently against a rubric you disagree with rather than substituting your own clinical preference. Expect questions about licensure status, jurisdiction, and how many adolescent risk assessments you have personally conducted.

Logistics

Fully remote and asynchronous. Most contributors pick up batched work in blocks of 5–20 hours per week; some projects run calibration sessions on a fixed call. Content is emotionally heavy by design, and projects usually ask you to confirm you can work with detailed self-harm and suicide material repeatedly. Pay is per hour worked on task, with observed rates clustering toward the upper end for independently licensed clinicians with emergency or crisis-line backgrounds.