Last updated: 2026-09-30
Local AI for Marking Support: A Live Build, and Why It Isn't Ready Yet
This page describes a real, working piece of software: a pipeline that runs locally hosted language models over substantial student project reports and produces per-criterion marks, feedback, and a flagged summary for a human marker to read. It is not a description of a plan. It has been built, run, broken, and fixed several times over. Two things about its current status matter more than anything else in the sections below. First, every model it uses runs on local hardware; nothing leaves the machine, at any stage, for any reason. Second, it has so far only ever been run against archived project reports from roughly two years ago, work a human marker had already finished grading and moderating at the time. It has never touched a live submission, never produced a mark a student has seen, and never fed into an actual grade.local does not mean safe. logs, screenshots, backups
The aim behind the project is specific, and worth stating before anything else: to make marking more consistent from one report to the next and lighter on time, without taking away the reason the job is worth doing in the first place. A second, independent read of a report against the same rubric line is one of the more reliable ways to catch a marker's own blind spot, and the two-model cross-check described below exists to make that second read cheap enough to get every time, not to remove the first one. Final-year projects regularly turn up genuinely interesting work, and reading it closely is one of the better parts of the job — a tool built to flag inconsistency and draft the mechanical parts of feedback is meant to leave more of that time for the reading itself, not to replace the reading with a machine's summary of it.
The University of Reading's guidance on AI in assessment and feedback sets the bar this kind of tool has to clear before it can do anything more than that: "GenAI must not determine marks" and "GenAI must not independently produce evaluative feedback," with marks and evaluative decisions remaining "with appropriately qualified staff" throughout [1]. Bounded assistance is only described as appropriate "after human academic judgement has been made and where all outputs are fully reviewed" [1]. This project is written up here specifically because it is still short of that bar, not because it has reached it, and any move from private experimentation toward real assistive use would go through the university's own approval route for higher-risk proposals before it went near a live cohort.
Why Nothing Here Touches a Cloud Model FoundationalKnowledge that endures for decades — core principles
Two separate obligations point at the same architectural decision. Student work carries personal data, which is the obvious reason it can't be pasted into a commercial AI service. It also carries something a name redaction doesn't remove: the student's own intellectual property, in the form of their actual code, their actual experimental design, their actual unpublished conclusions. Anonymising a report strips the names out of it; it does not turn someone's original project into public material safe to hand to a third party. Reading's guidance draws this line explicitly and without exceptions for scale or intent: "student work must not be uploaded into GenAI tools or AI detection tools under the current interim approach" [1].
That rule has to bind the whole project, not just a hypothetical future live-marking step. A design that kept student work local for marking but sent it to a cloud model "just for testing" or "just for calibration" would be protecting exactly nothing — the same real reports, the same real student output, cross a network connection either way. So every stage of this build, including the messiest parts described below (comparing models against each other, checking whether a prompt change helped or hurt, diagnosing why a mark looked wrong), was built to run entirely against locally hosted models: an 8B-parameter instruction-tuned model and a larger quantised model, both served through Ollama on ordinary local hardware, with nothing else in the loop. Anonymisation still matters as a second, independent layer of care — it limits what a human reader of the intermediate files can identify — but it was never the control doing the actual work of keeping the material contained. Staying local was.quantisation trades precision for speed
What "Bounded Assistance" Would Need to Look Like FoundationalKnowledge that endures for decades — core principles
Reading's policy gives a fairly precise shape to design against, even for a tool still in the archived-data testing stage. Assistance is only in scope after "human academic judgement has been made," which is a stronger condition than a human simply checking the AI's output afterward — it means a human's own view of the work has to already exist before the tool runs at all. The pipeline enforces this literally: it will not process a report unless a file recording a prior human read of that report is already present on disk. No file, no run. The tool's own output is framed throughout, in its interface and its written documentation, as supporting evidence for a marker rather than a submitted mark, and every field it produces — summary, strengths, per-criterion mark, justification — is meant to be read in full before it informs anything, matching the policy's "all outputs are fully reviewed" condition. The remaining piece, transparency to students about what role the tool played and what it didn't [1], isn't yet a live question, since no student's report has been through this pipeline while their grade was being decided — but it's a condition any future step toward real use would have to satisfy, not an optional extra.
What Was Built Ephemeral / ToolingKnowledge that evolves in months to a year — check for updates
The working parts of the pipeline, as it stands, are unglamorous and mostly exist because an earlier, simpler version broke in a specific, observed way. Two independently run local models mark the same report against the same rubric, so a human reviewer sees where they agree and where they don't rather than trusting a single model's say-so. Marks and feedback are returned as schema-constrained structured output rather than parsed out of free text [2], because the free-text version had a habit of returning an unparseable mark and silently recording it as zero — a completed criterion could not be told apart from a genuinely failed one until the raw text was checked by hand. Reports too long for a model's context window are split into pieces, assessed piece by piece, and reconciled in a final pass rather than truncated or skipped. Every completed piece of a run is written to disk as it finishes, after an early run hung for half an hour on the last stage of ten and lost everything that had already been computed. And when the two models disagree, the disagreement that gets flagged for a human is the spread on the report's overall total, not a raw percentage difference on any one criterion — most of the individual criteria here are marked out of five or ten, where a one-mark difference between two models is normal noise, not a signal worth a reviewer's time.
Where It Broke, First Direction Ephemeral / ToolingKnowledge that evolves in months to a year — check for updates
The first serious problem showed up as over-generosity. One archived report's discussion of ethical considerations consisted of a standard participant consent-to-share declaration and nothing else — no discussion of the project's own ethical dimensions at all. The model still returned a middling mark for it, and its own written justification described the section as missing almost everything the criterion asked for while somehow not placing it in the lowest band. A second case was more consequential: a report later flagged through the university's ordinary moderation process for suspected academic misconduct received competent marks for its methodology and implementation sections from the same model, despite those sections reading as fluent, textbook-style descriptions of what a good methodology would involve without ever naming the actual dataset, tools, or decisions behind the student's own work. The model was rewarding the presence of the right vocabulary and the right section headings, not evidence that the underlying work had actually happened.
The Fix, and the Overcorrection Ephemeral / ToolingKnowledge that evolves in months to a year — check for updates
The direct fix was to add explicit instructions to the marking prompt: content that is absent, generic, or purely administrative goes in the lowest band regardless of how complete it reads, and a Good or Outstanding mark requires specific, checkable evidence — named datasets, named tools, actual parameters, actual decisions — not just correct terminology covering the right topics. Tested directly against both cases above, this worked. The consent-declaration section and the misconduct-flagged report's methodology and implementation both dropped to marks consistent with what they actually contained.
Run across a larger batch of archived reports, the same instructions overshot. Literature-review sections landed at the exact floor of the lowest band in six of seventeen reports; results sections did the same in seven. The clustering on one identical value, rather than a spread of low marks reflecting different degrees of weakness, was itself a sign something had gone wrong beyond ordinary harsh marking. Checked directly against report content, one methodology section dismissed as "generic" and "lacking specific details" in fact named a specific, real, correctly cited open-weights model and gave concrete reasons for choosing it over the alternatives it had been evaluated against. The instruction built to stop the model crediting vague language for real work had started penalising some real work for sounding, in places, like the vague language it was written to catch.the prompt is overfitting to examples
A Lesson About How Small Models Take Instructions Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
An attempt to make the same prompt shorter produced a clean, decisive result. The instructions above are wordy: repeated phrasing, a concrete worked example, capitalised emphasis on the words doing the real work. Trimming them down to a few plain sentences that kept every point but dropped the repetition and the example, tested against the misconduct-flagged report again, reverted the fix completely — the same marks that had correctly dropped came straight back up. A version that kept the example and the emphasis but cut some of the restated phrasing landed in between, catching one criterion correctly and only partly catching the other. The repetition and the worked example weren't padding; for this model, they were doing measurable work that the same information stated once, plainly, didn't do.more tokens : more attention to the rule
That sits next to a finding from independent research on using language models as judges: alongside position and self-enhancement bias, that research documents a measurable verbosity bias, where a judging model tends to favour longer, more elaborate answers from whatever it's evaluating [3]. That result is about the length of the material being judged; the one described here is about the length of the instruction doing the judging, which is a different claim connecting the same underlying model behaviour to a different part of the pipeline — worth naming as this project's own extension of that finding, not something the original research showed.
The Real Problem Is Calibration, Not One Rule FoundationalKnowledge that endures for decades — core principles
Underneath both failures sits a problem that no single prompt edit closed. A separate, still-open issue in this project is the opposite of the one above: genuinely strong reports scoring well above a mark independently arrived at through the university's own moderation process, apparently because the model doesn't distinguish "competently done" from "exceptional" reliably. Two attempts to fix that by adding a third instruction to the same prompt each made results worse rather than better — one produced a mark whose own justification contradicted itself in the same paragraph, the other quietly reopened the over-generosity bug the first fix had already closed. Tightening the prompt against one failure mode kept loosening it against another, in a pattern that didn't converge as more instructions were added. That is a different kind of problem from any individual bug in this write-up. It suggests the model isn't holding a stable notion of how much evidence is enough for a given mark; it's responding to surface cues in the prompt's own wording, and those cues shift the outcome in an 8B-class model under fast, non-reasoning generation more than they should.
Why Not Fine-Tune the Model? Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
An obvious next question is why the fix for a calibration problem was another round of prompt wording rather than fine-tuning the model directly on a set of past reports and their moderated marks. That option was considered and set aside, and the reason isn't effort — it's that nowhere near enough suitable data exists yet.
The rubric behind this pipeline has ten criteria, each marked against four grade bands, which is forty distinct combinations a fine-tuned model would need to learn to recognise reliably. Final-year project topics are chosen by the student, not assigned from a fixed list, and the reports used as examples on this page alone span a marketing-analytics dashboard, a machine-learning diagnosis model, and an LLM routing framework — three projects sharing almost no vocabulary. Teaching a model to spot genuine evidence versus generic description as a topic-independent skill, rather than memorising what a Good mark happens to look like in whichever topics were well represented, means each of those forty combinations needs real examples spread across many different topics, not a handful drawn from one or two. A deliberately generous estimate — a few dozen well-distributed examples per combination — puts a workable floor somewhere in the low thousands of independently marked, moderated reports. A single module producing on the order of a hundred reports a year would need the better part of a decade of consistently rubric-marked archives to reach that floor. The handful of anonymised reports available for this project, low dozens at most, fall well short of it.
Short of that floor, fine-tuning carries a real risk of making things worse, not better. A model trained on a data set skewed toward whichever topics happened to be available would plausibly learn to reward the phrasing and structure of those specific topics, rather than the general evidence-versus-description judgement this page has been chasing throughout — the opposite of what fine-tuning is meant to buy, and a worse outcome than a careful prompt applied to an untouched general-purpose model. Understanding Large Language Models covers what training data actually does to a model's behaviour, and why a narrow, skewed data set produces a narrow, skewed model. Combined with the ongoing cost of curating a clean training set and re-tuning it every time the rubric's wording changes, against a prompt edit that costs an afternoon, fine-tuning isn't judged cost-effective at this project's current scale, whatever its merits once — or if — enough marked archives exist.
What Local Models Are Actually Good For, Right Now Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
None of this makes the pipeline useless — it changes what it's honestly good for. Structured output means a completed assessment and a genuinely missing one are never confused with each other. Partial-result persistence means an interrupted run costs minutes, not hours. A full record of exactly what each model was shown and exactly what it returned makes every mark traceable back to its source, which a rushed human marker's private impression never is. And running two independent local models over the same report and flagging where their totals genuinely diverge is a real, useful signal — not because either model's mark is individually trustworthy, but because a wide disagreement between them reliably points a reviewer at a report worth a closer look, which is a different and much more modest claim than "this mark is correct."
What Happens Next Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
Two directions look more promising than another round of prompt instructions stacked on top of the ones already there: showing the model a small number of worked examples of genuinely evidenced versus genuinely generic sections rather than describing the distinction in the abstract, or splitting evidence-checking into its own separate pass instead of asking one prompt to weigh evidence and assign a calibrated mark at the same time. Either is worth testing properly before this tool goes anywhere near the "bounded assistance" role Reading's policy allows for, let alone a live cohort — and that step, when it comes, goes through the university's own sign-off process first, not through a decision made inside this project alone. Local models plausibly can do this job one day without a single student's work ever leaving the building it was written in. On the evidence collected so far, that day hasn't arrived.
One Person's Attempt, Not a Verdict FoundationalKnowledge that endures for decades — core principles
Everything on this page is the work of one person, testing in spare time on ordinary local hardware, without a research grant, a dedicated evaluation team, or a systematic search across prompts and model settings. Each fix described above was checked by hand against a small number of real cases, one change at a time — a reasonable way to catch an obvious bug, and a poor substitute for a controlled, statistically powered evaluation run by a properly resourced team. The mixed results reported here are evidence about what one relatively quick, manual attempt achieved. They are not a ceiling on what local models can do for this task. Someone with more time, a larger validated archive of moderated reports, or a more systematic approach to prompt or few-shot design could plausibly do better than this account suggests. The conclusion this page draws is "not ready yet, on this evidence" — not "cannot be done."
Related Topics
- Project Marking as BDD — turning the same vague rubric language this pipeline struggles with into concrete, checkable scenarios for human markers.
- Evaluating Detailed Rubrics — the evidence on structured criteria and cross-checking between independent markers that this pipeline's two-model design is a machine analogue of.
- Academic Misconduct and GenAI — the wider context for the misconduct-flagged case this page uses as a real test case.
- Jev and "System One" Models — calibration as a measurable, task-specific property of a model's stated confidence, the same idea underlying this page's unresolved top-end problem.
- Legal Framework in Computing — the data-protection landscape behind the local-only constraint described above.
References
- University of Reading, Centre for Quality Support and Development. Artificial Intelligence: Assessment and Feedback. https://www.reading.ac.uk/cqsd/artificial-intelligence/assessment_and_feedback
- Ollama. Structured Outputs. https://docs.ollama.com/capabilities/structured-outputs
- Zheng, L., Chiang, W.-L., Sheng, Y., et al. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. Advances in Neural Information Processing Systems (NeurIPS) 36, Datasets and Benchmarks Track. arXiv:2306.05685