Independent specialist
Prompt, Response and Search Relevance Evaluation
Evaluation needs more than a quick opinion. I apply repeatable criteria to prompts, model outputs and search results, document edge cases and help teams understand where quality breaks down.
What this service can include
- ✓ Prompt clarity and instruction-following review
- ✓ Response quality and factuality scoring
- ✓ Search-result relevance evaluation
- ✓ Side-by-side model comparison
- ✓ Error categorization and written feedback
Who it is for
Useful for AI product teams, search teams and dataset owners who need dependable qualitative judgment before scaling.
How the work moves forward
- 01Task and rubric review
- 02Calibration examples
- 03Structured evaluation
- 04Quality-control checks
- 05Pattern summary and recommendations
Frequently asked questions
How do you keep evaluations consistent?
I use the supplied rubric, calibration examples, decision notes and recurring quality checks.
Can you compare two model responses?
Yes. Pairwise evaluation can cover correctness, relevance, clarity, safety and task-specific requirements.
Do you create evaluation guidelines?
I can help improve a rubric or document edge cases discovered during a pilot.
Related capabilities
Explore connected services.
Have a project in mind?