What the work looks like
You run structured evaluation tasks that mirror ordinary analyst days — investigate an anomaly in a metric, produce a KPI report, refresh a stale dataset — except the work is done through an AI assistant wired into a cloud data warehouse. Your job is to be the adversary who checks it. When the assistant reports Q3 net revenue retention at 112%, you write your own SQL against the seeded Snowflake data and find out whether the join fanned out, whether the filter silently dropped nulls, and whether the time window matched the definition the prompt implied. Findings go into rubric scores plus written notes that a model team can act on.
There is a real infrastructure component. You maintain and reset seeded datasets so each scenario starts from a known correct answer state, manage warehouse roles, grants and access controls so environments are secure and repeatable, and configure and document connectivity and authentication across analytics and SaaS platforms. Expect to hit undocumented product behavior and be asked to write it up clearly enough that someone else can reproduce it.
What the screen is looking for
- Advanced SQL you can defend under follow-up, not SQL you have seen — aggregation traps, join cardinality, window functions, timezone and boundary handling.
- Concrete Snowflake fluency: warehouses, roles and grants, query history as a debugging tool.
- A reconciliation instinct: the habit of proving a reported number rather than sanity-checking it.
- Fluency in KPI definitions as leadership and external stakeholders actually use them.
- Calibrated judgment — consistent scoring against a rubric, and the ability to explain a score to a peer who scored differently.
No AI or ML background is required. Prior use of AI assistants for analytics helps mainly because it makes you appropriately skeptical of generated SQL.
Logistics
Contractor engagement, fully remote, on a confidential client project. Work is largely asynchronous and task-based, with periodic live calibration sessions with peers to keep grading consistent — expect at least some overlap with the client's working hours for those. Observed pay for this listing is $40–60/hr, which typically tracks depth of Snowflake and reconciliation experience; rates are as observed and not guaranteed. Volume on evaluation projects fluctuates, so treat it as substantial part-time to full-time contract work rather than a fixed schedule.