Independent specialist

Prompt, Response and Search Relevance Evaluation

Evaluation needs more than a quick opinion. I apply repeatable criteria to prompts, model outputs and search results, document edge cases and help teams understand where quality breaks down.

What this service can include

  • ✓ Prompt clarity and instruction-following review
  • ✓ Response quality and factuality scoring
  • ✓ Search-result relevance evaluation
  • ✓ Side-by-side model comparison
  • ✓ Error categorization and written feedback

Who it is for

Useful for AI product teams, search teams and dataset owners who need dependable qualitative judgment before scaling.

How the work moves forward

  1. 01Task and rubric review
  2. 02Calibration examples
  3. 03Structured evaluation
  4. 04Quality-control checks
  5. 05Pattern summary and recommendations

Frequently asked questions

How do you keep evaluations consistent?

I use the supplied rubric, calibration examples, decision notes and recurring quality checks.

Can you compare two model responses?

Yes. Pairwise evaluation can cover correctness, relevance, clarity, safety and task-specific requirements.

Do you create evaluation guidelines?

I can help improve a rubric or document edge cases discovered during a pilot.

Related capabilities

Explore connected services.

Have a project in mind?

Work directly with the specialist doing the work.

Tell me about your project ↗