What the work actually involves

You will spend most of your time reading model outputs and judging them against clinical reality. That means reviewing simulated patient interactions where an AI plays a clinician or responds to a service member in distress, marking where the reasoning breaks — a missed suicide risk escalation, a PTSD criterion applied without the required duration or functional impairment, a TBI presentation misread as purely psychological, a response that talks to a Marine like a civilian outpatient. Alongside evaluation, you will write: original prompts and scenarios drawn from real presentations (post-deployment reintegration, MST, moral injury, chronic pain plus opioid history, VA disability-rating friction), reference answers other reviewers are scored against, and short structured rationales explaining why a given output fails. Some projects add rubric calibration calls or consultative sessions with the customer's research team.

What the platform screens for

micro1 runs an AI-led interview before any human contact. It verifies the credential first — completed residency and board certification for psychiatrists, conferred PhD for psychologists — then probes whether your military/veteran experience is direct patient care rather than adjacent policy, research, or administrative work. Expect follow-ups that push on specifics: which setting (VA medical center, MTF, Vet Center, embedded behavioral health, community provider with a heavy veteran panel), roughly what caseload, which conditions, over what period. The screen also tests evaluation judgment — whether you can say precisely what is wrong with a plausible-sounding clinical response and rank two flawed answers against each other — and whether your critique holds up when the interviewer pushes back. Prior AI experience is not required and not scored.

Logistics

  • Remote, contractor, invoiced hourly against approved work.
  • Largely asynchronous; most reviewers work in blocks they schedule themselves, with occasional live calibration sessions in US business hours.
  • Volume fluctuates by project phase — expect uneven weeks rather than a steady load, and be honest about the minimum hours you can hold.
  • Non-English proficiency (Japanese, French, Tagalog, Malay, Lithuanian, Russian, Georgian, Polish, Marathi, Bengali, Urdu, Vietnamese, Swahili, Mandarin, Cantonese, Indonesian, Turkish, Gujarati, Bulgarian) is prioritized and often routed to separate, better-paid queues.
  • Written output is the deliverable. Reviewers whose rationales are vague get fewer batches regardless of clinical seniority.