What the work actually looks like
Projects vary, but BI work on AI evaluation platforms usually falls into a few shapes. You might write realistic analytics prompts — a messy schema, an ambiguous stakeholder request, a metric definition that conflicts across two source tables — and then author the reference answer yourself. You might grade model output: checking whether a generated SQL query actually returns what the question asked, whether a proposed star schema handles the stated grain, whether a dashboard recommendation reflects how a real business user would consume it. Other assignments lean toward rubric writing and adjudication, where you resolve disagreements between two earlier graders and explain your reasoning in writing.
The hard part is rarely the SQL. It's articulating why an answer is wrong in a way that transfers — distinguishing a query that's syntactically valid but silently wrong from one that's merely inelegant, or noticing that a model confidently invented a join key. Labs pay for that judgment, and written explanations are usually weighted as heavily as the pass/fail call.
What the platform screens for
Mercor's intake is resume verification plus an AI-conducted interview. The interview probes specifics: which warehouse you worked in, how large the tables were, what a particular metric definition actually meant at your company, how you handled a modeling decision you later regretted. Vague answers get follow-ups until they aren't vague. Expect to name tools, describe schemas, and defend a technical choice under pressure. Passing the interview makes you eligible; it does not create an assignment.
Logistics
- Fully remote, largely asynchronous, with occasional project-specific calibration sessions
- Typical commitment 15–30 hrs/week when matched; matching is rolling and there can be gaps between projects
- Pay band observed at $70–120/hr — actual rates are set per project by the client lab, not guaranteed by the listing
- Contract work, self-scheduled within project deadlines