What the work actually involves

You are the materials authority on a team building training and evaluation data for frontier models. Day to day, that means three overlapping streams. First, review: reading model outputs and contributor-authored tasks and deciding whether the reasoning would survive scrutiny from a reviewer in your subfield — catching unsupported structure–property claims, hand-waved processing conditions, characterization interpretations that don't follow from the data shown, and answers that are fluent but wrong. Second, authoring: writing instruction specs that define what correct looks like, and producing golden solutions to materials problems that reflect how the work is genuinely done rather than how a textbook presents it. Third, benchmarks: designing evaluation sets hard enough to discriminate between a model that knows the vocabulary and one that can actually reason, and helping the research team build materials-specific tools and skills.

A recurring demand is calibration. Much of what makes a senior materials scientist good is tacit — the sense that a reported capacity fade curve looks fabricated, or that a proposed alloy heat treatment would coarsen the wrong phase. This role asks you to convert that into explicit, teachable criteria that a researcher outside your subfield, or a grader, can apply consistently. You will work alongside client researchers and experts in adjacent fields to keep standards aligned.

What the screen looks for

  • A named specialization, defended under follow-up. Energy storage, semiconductors and electronic materials, polymers and soft matter, structural alloys and metallurgy, characterization and microscopy, or computational materials. Generalist breadth is not what this listing is buying.
  • Post-PhD depth with ownership. 4+ years of research or industrial R&D at a university, national lab, or industrial research organization, with clear progression to senior level and real ownership of research direction. Coursework does not count.
  • Evidence you can write. Instruction specs and review feedback are the deliverables; the screen weighs how precisely you express a technical judgment in writing.
  • Genuine LLM use in your own work, including the ability to describe a case where a model produced something plausible and wrong in your field.

Logistics

This is not remote. It is hybrid, based in the Bay Area, requiring on-site presence with the client's team multiple days each week when needed. If you are not already in the Bay Area you must relocate at your own cost before the engagement starts; no relocation assistance is offered. The commitment is 40 hours per week for an initial six-month engagement, as W-2 employment with Cincinnatus LLC acting as employer of record, with client-issued accounts and equipment and work performed inside the client's tools. Observed pay for comparable Mercor domain-expert placements has been in the $70–110/hr range; nothing here guarantees a specific rate.