What the work involves

You are writing the problems, not solving someone else's. A typical task starts with a realistic drug discovery scenario — filter a ChEMBL activity set for a target, flag assay inconsistencies, compute descriptors and cluster a series, triage a virtual screening hit list, or reason through an ADMET liability in a lead series — and ends with a runnable environment plus tests that verify a model actually got the science right. Expect to spend real time in Python and RDKit, in a terminal, and in Docker, then write the accompanying rationale explaining why a wrong answer is wrong.

The second half of the work is evaluation. You review AI-generated code, analyses, and chemistry rationales for scientific rigor: does the model understand that an IC50 from a different assay format isn't comparable, that a SMILES standardization step matters, that a QSAR model trained on that split is leaking? Feedback is written, specific, and cites the chemistry or biology — not "looks plausible."

What the screening looks for

  • Verifiable domain depth. Named targets, series, assays, and toolkits you have personally worked with, and what went wrong.
  • Actual coding ability. Python beyond a notebook script: packaging, CLI tools, pytest, Git, Docker. This is a genuine filter, not a preference.
  • Task-design judgment. Whether you can write a problem that is unambiguous, solvable, and has a defensible ground truth.
  • Calibration. Willingness to say "this is underdetermined" instead of grading confidently on a bad prompt.

Logistics

Fully remote, contractor, asynchronous. Contributors typically commit 10–20 hours per week with occasional scoped bursts; there are no fixed shifts, but tasks carry deadlines and rework requests. The $80–110/hr range reflects rates observed on this listing and varies with credentials, coding depth, and task type — treat it as observed, not guaranteed. Advanced degrees (PhD, MSc, PharmD) are strongly valued but industry track record can substitute.