What the work involves
This is a short, concrete evaluation task rather than an ongoing contract. You receive a set of prompts and scenarios, dial the client's conversational AI agent from a line that can reach an EU, UK, or US number, and work through each scenario in Arabic as a realistic customer would — including the awkward turns where the agent is most likely to break. You then complete ten QA reviews of calls, scoring accuracy, tone, clarity, and naturalness, and writing short free-text feedback on conversation flow and overall user experience.
The judgment being paid for is linguistic, not technical. A synthesised Arabic voice can be grammatically correct and still land wrong: register that is too formal for a service call, MSA where a dialectal reply was expected, mispronounced proper nouns, prosody that makes a question sound like a statement. The useful reviewer names the specific turn where the problem occurred and says what a natural speaker would have said instead.
What the screen looks for
- Native-level Arabic, with an ability to describe your own dialect and how far it diverges from what the agent produces
- Prior annotation or QA experience — the platform expects you to already know what a rubric is and how to apply one consistently across ten items
- Working call setup: a phone or VoIP line that can place international calls to EU/UK/US numbers with clean audio, plus a quiet environment
- Genuine immediate availability — this task is staffed for a start date of now, not next week
Logistics
Fully remote and largely async, though the calls themselves must be placed live during hours when the agent is available. Total commitment is 1–2 hours. Pay in this band has been observed at $16–18/hr for similar Mercor evaluation tasks; it is not guaranteed, and short tasks like this are sometimes paid as a fixed amount for completed batches. Expect the possibility of repeat invitations if your written feedback is specific and your scores are internally consistent.