The work

This is task authorship for a large-scale scientific benchmark, not annotation. You pick a domain you genuinely know — Bayesian inference in Stan, IRT in mirt or TAM, spatial point processes in spatstat, state-space models in KFAS or MARSS, survival analysis in flexsurv, ODE inference in pomp or deSolve, high-precision linear algebra in Rmpfr, or the Matlab/Scilab equivalents — and build problems that require a model to actually run the software correctly rather than pattern-match a plausible answer.

Two broad problem types recur. The first is fully specified and reproducible: a defined setup with a single correct numerical answer that can only be reached through a multi-step workflow. The second is interactive and harder: the model must plan a sequence of queries or simulated experiments to recover information that is not directly observable, reason from partial results, and narrow the hypothesis space efficiently. Both types require you to write the setup, an oracle function, and a validator in Python, then run the problem against current frontier models and revise until the difficulty is right — too easy and it gets cut, too ambiguous and it gets cut.

What the screen looks for

  • Verifiable depth in a specific package. Expect follow-ups on convergence diagnostics, identifiability, parameterization choices, numerical stability, and where your tool of choice quietly gives wrong answers.
  • Evidence, not claims. Publications, open-source contributions, or professional work that show you have used these packages on real research problems.
  • Problem-design instinct. Can you distinguish a hard problem from a merely tedious one, and can you specify a task tightly enough that a validator can judge it?
  • Python fluency and terminal comfort. You will write oracles and validators and work inside Linux-based remote compute sandboxes.

Logistics

Fully remote and asynchronous, with roughly 15–20 hours per week expected. Credentials required are MS or PhD in statistics, applied mathematics, or a closely related quantitative field — PhD preferred, or MS with 10+ years of relevant experience. The observed pay band is $70–90/hr; rates on Mercor vary by domain, assessed depth, and project stage, and are set at offer rather than guaranteed by the listing.