What the work involves

You receive artifacts an AI model produced in response to a legal prompt: a trademark clearance summary, a fair-use analysis, a licensing term sheet, a patent landscape spreadsheet, or a client-facing deck on brand enforcement strategy. Your job is to grade them the way you would review a junior associate's draft — but against a written rubric, with scores and written justification. That means checking whether the likelihood-of-confusion factors are applied to the right circuit's test, whether a cited case actually says what the output claims, whether the copyright term calculation accounts for the correct authorship and publication facts, and whether the deck's formatting and structure would survive a partner review.

Most tasks ask for a numeric score plus a short written rationale tied to specific defects. Strong evaluators are specific and traceable: not "the analysis is weak" but "the output treats the DuPont factors as exhaustive and never addresses the applicant's own prior registration, which is dispositive here." You will also flag presentation-level problems — mislabeled exhibits, broken slide hierarchies, spreadsheet formulas that don't compute — because these tasks measure whether the artifact is usable, not just whether the law is right.

What the platform screens for

  • Verifiable practice history. Expect follow-ups on where you practiced, what kinds of matters, and whether your IP work was prosecution, litigation, transactional, or in-house counseling. Vague answers get probed.
  • Depth under pressure. The screen tends to pick one thing you claim and drill: a doctrine, a statutory framework, a procedural posture. Depth on your actual subspecialty beats broad familiarity.
  • Evaluation judgment. Can you distinguish a confidently-worded but legally wrong answer from a hedged but correct one, and rank defects by severity rather than listing everything you noticed?
  • Deck and document fluency. Slides matter here more than in most legal roles; the requirement is real.

Logistics

Remote and asynchronous. Work arrives in batches tied to project demand, so weekly volume fluctuates — some evaluators see steady 15–20 hour weeks, others see gaps. Pay is hourly at the observed $80–120 range, generally set by seniority and specialization at onboarding rather than negotiated per task. There is usually a short calibration phase where your scores are compared against reviewer consensus before you move to production volume.