What the work involves

You will spend most of your time reading model outputs and candidate training data and deciding whether the developmental claims in them hold up. That means judging whether an AI's description of a four-year-old's theory-of-mind capacity is defensible, whether a suggested parenting response is age-appropriate, whether a scenario involving a child's language production is plausible for that age and linguistic background, and whether a piece of advice is culturally portable or quietly assumes a Western middle-class family structure. Alongside evaluation, contributors write scenarios and test cases — realistic situations involving children's learning, behavior, and distress — and draft or revise the rubrics and guidelines that other annotators apply.

  • Rating and ranking model responses on developmental accuracy, safety, and cultural fit
  • Annotating datasets with age-band, construct, and confidence labels
  • Authoring adversarial or edge-case prompts where models are likely to overgeneralize from norms
  • Writing rationales that a non-psychologist reviewer can follow and audit

What the platform screens for

micro1's screening is AI-led and conversational. It probes whether your doctorate is real and recent enough to be usable, whether you can distinguish an established developmental finding from a contested or replication-fragile one, and whether you can hold a position under follow-up without either collapsing or overclaiming. Expect to be pushed on specifics: which stage models you actually use, how you'd handle a normative claim that varies across cultures, what you'd do with a model output that is technically accurate but clinically unwise. Prior AI experience is genuinely not required, but comfort with structured rubrics and written rationale is.

Logistics

Contractor engagement, fully remote, asynchronous. Work arrives in batches; contributors typically commit to a stated weekly hour range rather than fixed shifts, and volume fluctuates with project phase. Rates in the $100–300/hr band have been observed on this listing and vary with credentials, language coverage, and task type — nothing is guaranteed. Proficiency in one of the prioritized languages (Japanese, French, Tagalog, Malay, Lithuanian, Russian, Georgian, Polish, Marathi, Bengali, Urdu, Vietnamese, Swahili, Mandarin, Cantonese, Indonesian, Turkish, Gujarati, Bulgarian) materially improves selection odds but is not required.