The work
You'll be handed frontier-model outputs on physics problems — classical mechanics, E&M, statistical mechanics, quantum, condensed matter, cosmology, or whatever your subfield happens to be — and asked to say precisely where the reasoning breaks. That means checking dimensional consistency, tracking sign and factor errors through a derivation, catching plausible-sounding but wrong limiting behavior, and distinguishing a genuine physical insight from a memorized result. A second stream of work is authoring: writing original problems hard enough that current models fail them, with a fully worked solution and a defensible reference answer.
Typical tasks include:
- Rubric-based grading of multi-step derivations, with written justification for each deduction
- Head-to-head preference comparisons between two model responses
- Authoring novel problems at qualifying-exam or research-adjacent difficulty
- Flagging cases where a model reaches the right number by an invalid route
What the platform screens for
micro1's screening is AI-led: an asynchronous conversational interview plus, usually, a technical exercise. It probes whether you can actually do physics under follow-up questioning rather than describe having done it. Expect to be asked to reason aloud about an unfamiliar setup, to name your dissertation topic and defend a methodological choice in it, and to explain an error you once made and how you found it. Vague answers get re-probed; specificity is the whole game. Written English clarity matters, because your grading rationale is the deliverable, not just your score.
Logistics
Fully remote and largely asynchronous, with work claimed from queues rather than scheduled shifts. Most contributors commit 10–20 hours weekly, though project surges open more; calibration onboarding and periodic reviewer audits are standard. Engagements are contractor-based and scoped per project, so volume fluctuates — treat it as a supplement to academic or industry work rather than a fixed salary. Rates in the $70–90/hr range have been observed for PhD-level physics work and vary by task type, subfield scarcity, and quality scores; nothing is guaranteed.