The evidence

Every extra question is another judgement someone has to make

GradeLogic™ drafts every one of them for you to confirm or overturn. Here is what that does to a stack.

  • Precision

    The last exam is graded as carefully as the first.

    The rubric does not soften at the bottom of the stack. There is no session for the model to get tired in.

  • Speed

    Finished while the red pen is still on page one.

    Upload the stack, walk away, come back to drafted scores and a log of everything the model was unsure about.

  • Cost

    Grader hours go back to teaching.

    The grader hours you were buying to get through a stack go back to office hours, or back to the department.

What the research already measured

  • Procedural, apply a known step1.82s
  • Factual, check recall against a key7.95s
  • Conceptual, judge whether they understood14.73s

Eight times the cost for the question actually worth asking. Judgement, not length, is what makes grading expensive.

More questions means more responses

Workload is students times responses times effort per response. A 438-student exam answering 36 questions each is 15,768 separate judgements before anyone reads a word.

Deeper questions cost more per response

Timed at 1.82 seconds for a procedural answer, 7.95 for factual recall and 14.73 for a conceptual one. Judgement, not length, is what makes marking expensive.

Long sessions change the grading itself

Shorter shifts scored more accurately and more consistently than long ones. In a smaller essay study, both scores and written feedback shifted as the stack went on.

Weegar and Idestam-Almquist, IJAIED 34 (2024). Ling, Mollaun and Xi, Language Testing 31 (2014). Mahshanian and Shahnazari, IJAL 10 (2020).

Run the same arithmetic on a real exam

Every exam here is one we ran. The hand-grading figure is an estimate from the published seconds per response above; everything labelled measured came off the clock.

Seventeen papers, six multiple-choice questions each, eighty-five scanned pages. Splitting and identity marking by hand, then the run went to Auto-Pilot and finished without a stoppage.

0%of each paper needs judgement. 6 of 6 answers match a key.

  • 17submissions
  • 85pages
  • 102responses
  • 6Multiple choice7.95s each, measured100% of the grading costsingle answer, between two and twelve options

13m 31s

Marking alone, estimated13m 31s
The instructor's own timed baselineNot measured yet
GradeLogic™, measured5:29
  • 102responses to judge, 6 from each of 17 papers
  • Not measured yetof the run needed the instructor. Attended time was not tracked on this run.
  • Not measured yetof scores kept. Not one answer on this run was checked against a key.

Seconds per response measured in Weegar and Idestam-Almquist, IJAIED 34 (2024).

Now watch one run against the red pen

17 papers on one wall clock. GradeLogic™ ran the time you see; the red pen is what the published rates say the same stack would have cost. Drag to any second.

5:29of 5:29, run complete

GradeLogic™results on screen
  • Exam splitting—
  • PII annotation—
  • Zone detection—
  • Solution inference—
  • Answer parsing—
  • Autograding—
  • Curving—

Per-phase timings were not recorded on this run. The 5:29 total is measured; the split across the seven steps is not.

Hand grading, estimated13m 31s

6 of 17 papers graded at the published rate

  • 17 vs 6papers finished by 5:29, by GradeLogic™ and by the published hand rate
  • Not measured yetcorrections. No score on this run has been compared against a key.
  • Not measured yetof instructor time. How much of the run ran unattended was not tracked.
  • No roster was linked, so student matching was skipped entirely.
  • Splitting and identity marking were done by hand. Everything from zone detection onwards ran unattended.
  • Every question was multiple choice. Multiple choice is the cheapest kind of exam to grade by hand, which makes this the least flattering comparison we could have published.
  • Not one answer on this run was checked against a key. Until that happens the speed above says nothing about whether the scores are right.

Running total, one published run

8m 02s

of grading time returned across 17 papers — the gap between what the published rates say the stack costs by hand and what GradeLogic™ measured.

  • Papers put through17
  • Responses judged102
  • Marks checked against a keyNot measured yet
  • Non-Coding Quiz 1 Section 28m 02s

Early access · waitlist open

Put one stack through it and read the log

Setup, splitting and review cost nothing. Credits are spent only when the AI grades.

Hand-grading times on this page are estimates derived from published research, never something we watched an instructor do. GradeLogic™ times are measured on real runs. No score produced by any run here has yet been checked against an answer key.