AI QA Trainer - LLM Evaluation
A USD 82K-131K (estimate) Senior-level Contract Full Time
Tasks
- Analyze model failures
- Assess tool usage accuracy
- Automate testing with Python and SQL
- Collaborate on evaluation metrics and dashboards
- Conduct robustness evaluations
- Create evaluation frameworks
- Design test plans
- Develop evaluation rubrics
- Document root cause analysis
- Evaluate large language models
- Execute regression suites
- Maintain evaluation documentation
- Perform red teaming
- Recommend prompt and guardrail improvements
- Run adversarial testing
- Set PASS/FAIL criteria
- Validate retrieval augmented generation outputs
- Verify grounding
Perks/Benefits
Skills/Tech-stack
Bias Assessment | Experiment tracking | Hallucination detection | LLM Evaluation | Language Models | Large Language Models | Machine Learning | Prompt engineering | Python | RAG | Red Teaming | Regression testing | Retrieval-Augmented Generation | Robustness Testing | SQL | Safety testing | Test automation
Education
Roles
AI | AI QA | AI QA Trainer | Engineer | Learning Engineer | Machine Learning Engineer | QA Trainer | Trainer
Related jobs
-
Alerting | Amazon Redshift | CI/CD | Cause analysis | DashboardsContinuous improvement | Cross-functional collaboration | Fully remote work | Opportunity for ownershipSenior-level Full TimeSaudi Arabia R4d ago
-
Senior AI Backend Engineer - Agent Evaluation & Quality USD 175K-250KAPI Design | CI/CD | Cost Optimization | Experimentation | Function CallingSenior-level Full TimeMakkah, Makkah Province, Saudi Arabia - … R6d ago
-
Senior-level Contract Full TimeRiyadh, Riyadh Province, Saudi Arabia - … R16d ago