What the work involves

You'll be handed problems that sit at the boundary of analytic statistical mechanics and numerical benchmarking: replicated random-bond Ising and Ashkin-Teller models, the Nishimori line, square-lattice self-duality, and the toric-code threshold. Depending on the assignment and your subfield, you'll work in one of three seats — Solver (produce a fully worked, defensible solution), Auditor (find the error in someone else's derivation, or confirm there isn't one), or Adjudicator (rule between conflicting submissions and write the reasoning that settles it). Deliverables are written artifacts: derivations with stated assumptions, critiques that point to a specific line rather than a vibe, and occasionally simulation code or numerical results for 4-state Potts and domain-wall free energy calculations.

The underlying purpose is AI training data. Model outputs on this material are frequently confident and wrong in ways that require someone who has actually done a quenched disorder average to catch — a misapplied replica limit, a duality mapping that quietly drops a boundary term, a noncontractible loop defect handled as if the topology were trivial. No prior AI or ML experience is required or expected.

What the screen looks for

  • Real specialization, not adjacency. Expect follow-ups that push past the first correct-sounding answer: where the Nishimori line comes from, why self-duality fixes the critical point on the square lattice, what breaks in the replica argument.
  • Written precision. Ambiguous notation and unstated assumptions are the main failure mode; the screen samples how you write, not just what you know.
  • Evaluation judgment. Given a plausible but flawed derivation, can you localize the error and articulate why it matters, rather than rewriting from scratch?
  • Numerical honesty. If you've run Monte Carlo on Potts models, you'll be asked about finite-size effects, equilibration, and what your error bars actually mean.

Logistics

Contractor engagement, fully remote and asynchronous. Work arrives in batches; most contributors take on a defined number of tasks per week rather than fixed hours, and volume fluctuates with the customer project's needs. Rates in the $80–160/hr band have been observed for this listing, with placement typically reflecting subfield fit and whether you're solving, auditing, or adjudicating — treat these as observed, not guaranteed. Expect an onboarding calibration round before paid work begins.