What the work involves

You are the tax authority checking an AI's homework. A typical task hands you a prompt — a shareholder basis question, a multi-state apportionment scenario, a §1031 exchange with boot, a Form 5471 filing threshold — plus one or more model-drafted answers. You judge whether the reasoning is correct, whether the citation to the Code, regs, or a revenue ruling actually says what the model claims, and whether the answer is safe to put in front of a taxpayer. When it fails, you rewrite it and explain what broke.

Other task types you should expect:

  • Preference ranking — two answers, both plausible, and you pick the better one with a written rationale that a reviewer can audit.
  • Prompt authoring — writing tax questions hard enough that a strong model gets them wrong, drawn from situations you have actually handled.
  • Rubric-based scoring — grading against dimensions like citation accuracy, treatment of the current tax year, and whether the model hedged when the law is genuinely unsettled.
  • Red-teaming — probing for confident answers about repealed provisions, expired TCJA sunsets, or state rules the model has conflated with federal.

What the screen looks for

micro1's screening is AI-led and conversational, and it follows up. It cares that your CPA licence or EA enrollment is current and verifiable, that your stated specialties survive a technical probe, and that you can explain why an answer is wrong rather than simply flagging it. Reviewers who write "incorrect" with no cited authority do not last on these projects. Fluency in US federal tax is assumed; state-level depth, international, partnership, or trust and estate work is where the higher end of the band tends to sit.

Logistics

Fully remote and almost entirely asynchronous — tasks land in a queue and you clear them against a deadline rather than a schedule. Most contributors commit 10–20 hours a week, and volume moves with client demand, so treat it as supplementary rather than a replacement for practice income. Expect an unpaid calibration exercise before paid work begins, and expect throughput to matter: consistently slow reviewers get fewer batches.