EducAItors is an AI-assisted evaluation platform designed to make grading more consistent, explainable, and auditable across instructors, courses, and institutions. My work focused on the Pre-Evaluation phase: the part of the system where instructors define assignments, structure rubrics, establish expectations, and prepare the logic that grading depends on.
The problem wasn't grading. It was everything that happened before grading began.
An instructor might receive a score of 73/100 without being able to trace exactly why. Students received generic feedback, while institutions had little evidence to explain or defend how an evaluation was reached.
The system was trying to make evaluation faster, but speed without clear criteria created another problem: trust.
I traced the problem upstream to assignment creation. Expectations were often unclear, rubrics inconsistent, and evaluation standards lived largely in the instructor's head.
The question became:
How do you make evaluation reliable before AI ever starts grading?
I redesigned the Pre-Evaluation experience around three stages: define the assignment, structure the evaluation, and validate before publishing. The system captured instructor intent through rubrics, learning-outcome mappings, calibration, and a final preview rather than asking AI to infer standards later.
The central design decision was to separate what the instructor decides from what the system can structure.
AI could suggest improvements, map questions to learning outcomes, identify rubric gaps, and organise evaluation logic. But every consequential decision remained visible, editable, and explicitly confirmed by the instructor.
AI should carry the structural load, not the responsibility.
Institutions use different rubric formats, scoring systems, and performance levels. I introduced a normalization layer that converted these variations into a consistent internal structure while allowing instructors to continue working in familiar formats. This created a common foundation for consistent evaluation without forcing every institution into the same visible rubric.
Traditional calibration depends on sample submissions, but that happens too late for a system trying to establish expectations before evaluation begins. I reframed calibration as intent capture, asking instructors to define what strong and weak answers should contain, what should be penalized, and what evidence would demonstrate quality.
Calibration could also be conditional: reused assignments with minor changes could inherit existing intent, while substantial changes triggered recalibration.
Make the standard explicit before asking the system to judge against it.
The Pre-Evaluation phase became more than an assignment setup flow. It became the foundation connecting assignment intent, rubric structure, learning outcomes, calibration, evidence, and downstream evaluation.
The system reduced setup effort from an estimated 3–4 hours to 45–60 minutes, while shifting instructor effort from repetitive structural work toward reviewing and refining evaluation logic.
EducAItors changed how I think about AI product design. The hardest problem wasn't deciding where to put AI. It was deciding where AI should stop. By automating structure while preserving human judgement, the product could become more efficient without making the instructor feel displaced.
Trust isn't created by automating the decision. It's created by making the decision understandable.