--- name: writing-assessments description: Write something that measures what students actually learned. Use when a teacher needs a quiz, a test, an exit ticket, a set of practice questions, or a rubric for a project or piece of writing. Starts from what the assessment is for, because a check to find out who needs reteaching and a test that produces a grade are different instruments, then builds items matched to the outcome with distractors drawn from real misconceptions, an answer key, and rubric levels that describe observable differences. For grades 1 through 8. --- # Writing assessments The quiz gets made on Thursday evening out of whatever the week turned out to cover. On Monday there is a column of numbers, and the teacher knows Ava got seven and Jonah got four, and nothing about what either should do next. Nothing in that is careless. Nobody decided what the instrument was for, and a quiz built to be both a check and a grade is reliably neither. The router (`using-gpa-skills`) has already screened for what does not belong in a planning conversation, drawn the line at diagnosis, and set how much to carry about any particular child. Rely on that. ## Where this starts, and where `planning-curriculum` stops `planning-curriculum` decides what evidence would show an outcome landed — a performance task, an exit ticket after lesson 3, a hinge question in lesson 6. It names the evidence and stops. This skill turns one of those names into a real object a class can sit down in front of. Two things follow. **Where there is no outcome yet, do not invent one here.** "We did the Great Lakes for three weeks" is a topic, and a quiz built on a topic can only ask what the children happen to remember. Ask the teacher for the sentence — what should they be able to do now that they could not before — or route to `planning-curriculum`, which sharpens it properly. **Do not rebuild that skill's alignment table.** Whether the unit taught what it assessed is its question, at unit grain. The mapping here is narrower and always present: one row per item, against the piece of the outcome it is evidence for. ## Why the method, and where to say so An assessment gets questioned — by a parent looking at a mark, by a colleague handed the rubric to co-mark, by the teacher herself in June. Name the reason in a line where it is carrying a decision, and nowhere else. - **Purpose decides the instrument** — Wiliam, *Embedded Formative Assessment* (2011), building on Black and Wiliam (1998). What makes an assessment formative is what happens to the result, not how short it is. - **Validity belongs to the interpretation, not the paper** — Messick (1989). His two failure names are the two below: a paper that reports on something it was never built to measure, or that samples too little of what it was. - **Distractors built from documented conceptions** — Philip Sadler (1998) built science instruments whose wrong options were the answers students actually give, so the choice a child makes identifies what she believes; the Force Concept Inventory (Hestenes et al., 1992) is built the same way. Note the surname collision: this is not the Sadler in the next line. - **Rubric levels describe, not judge** — Brookhart (2013). And D. Royce Sadler (1989): a child can only close a gap she can see, which is what a described level gives her and a labelled one does not. Jonsson and Svingby (2007) for why separate criteria mark more consistently than one overall impression; Andrade (2000; 2013) for what changes when the rubric reaches the child before the task rather than stapled to it afterwards. - **Retrieval and spacing** — Roediger and Karpicke (2006) for the testing effect, Rowland (2014) for its size across a large body of studies, Agarwal et al. (2012) for it holding up in real classrooms rather than laboratories, Dunlosky et al. (2013) for practice testing and spacing being the two that earned a high rating. Cepeda et al. (2006) for the gap. - **The verb ladder** — Bloom (1956), revised by Anderson and Krathwohl (2001). A check on the verbs in an outcome, and a planning heuristic rather than a measured fact about minds, so never defend an item by citing a level. ## Ask what it is for, before anything else One question, and it comes first: **is this to find out who needs reteaching, or to produce a mark that goes on the record?** | | To find out | To go on the record | | --- | --- | --- | | Length | short enough to fit inside a lesson with the reteaching | long enough to sample the whole outcome | | Coverage | one or two outcomes, at the point they are wobbling | everything the unit claimed, in proportion | | Distractors | the entire point of the exercise | still misconception-built, and now an ambiguous one costs a child | | Marking | fast, because you are reading a pattern rather than totalling | defensible to a second marker and to a parent | | Result | which children, which misconception, which lesson to redo | a mark, on whatever scale the school uses | | Stakes | none, and tell the class so | real, so the children know in advance what is coming | There is a third, rarely called an assessment: **a check run before the teaching.** Its score means nothing, and the pattern of wrong answers decides how the next two lessons go — which is why the hinge question in `planning-curriculum`'s worked unit sits before the lessons it informs. Results from one of these never become marks. What settles the purpose is what happens to the result, not the format; one quiz can be either. If the teacher says both — who is stuck *and* a grade this week — say plainly that these pull apart, and offer the two-part version: a short check on Tuesday that nobody marks, and the graded piece on Friday built out of it. ## Read what the school has already settled - `school/gpa-context.md#standards-framework` — what the items map to. - `school/gpa-context.md#grade-bands` — the school's own words for the grouping. - `school/gpa-context.md#report-card-cycle` — whether this mark reaches a report card, and in what form: letter grades, standards-based marks, or comments. **Where any of those is blank, ask.** And the harder version of the same rule: **do not invent the marking scale.** Not a pass mark, not a percentage that counts as meeting a standard, not what a level 3 is worth, not a total the paper is out of. Those are the school's, and an invented one is the most dangerous thing this skill can produce, because it looks official — a teacher will defend it to a parent believing the school set it. Where nobody can answer, that degrades rather than stops: build the instrument, describe the levels, leave the totals and cut-offs off, and say in a line that you did and why, so the gap stays visible rather than filled. ## Gather the rest in one message Five things, asked together: 1. **The outcome, in her words, exactly.** Do not tidy the verb. The verb is what the next section runs on. 2. **The grade.** Not optional. 3. **How long the children get**, and when it runs relative to the teaching. 4. **What has actually been taught** — taught, not planned. This is the validity question arriving early, and it is cheaper to ask than to discover. 5. **What she already knows they get wrong.** Worth more than the other four put together. An observed misconception is a distractor you do not have to guess. If five does not come back, build anyway; if one or two are missing, ask again. ## Match the item to the verb The verb in the outcome says what would count as evidence. An outcome that says *explain why* is not assessed by an item that asks *which*: a child can pick the right option from four and be unable to say a word about why it is right, and the mark she gets says she can. | The outcome says | Evidence is | What does not count | | --- | --- | --- | | name, list, label | the term produced or picked out | — | | describe | what happens, in order, in her words | naming the parts | | explain why | the cause, stated by the child | any selected-response item, alone | | compare | two things and the basis she compared them on | one of them described well | | apply, use, solve | a case she has not been walked through | the numbers from Tuesday's example | | decide, justify | the choice and the reason for it | the choice | Two rules fall out of that table. **Selected-response items cannot carry an explaining outcome on their own.** They are good at the pieces an explanation is built from, and worth having for it — but somewhere in the assessment the child has to produce the explanation herself, written or spoken. Ten multiple-choice items on an "explain why" outcome is a well-made assessment of a different outcome. **Show the mapping, always.** Every assessment goes back with a line per item: the item, the piece of the outcome it is evidence for, and what it asks the child to do. Four minutes, and it catches the two things nothing else catches — two items that are secretly the same item, and a piece of the outcome with no item against it. For the second, add one or stop claiming it is covered. ## Distractors are the diagnostic instrument In a multiple-choice item, the only information produced is which wrong answer a child chose. A distractor nobody picks tells you nothing. A distractor nine children pick tells you what Monday is for. So every wrong option should be **the answer a child arrives at by making one specific mistake you can name** — and the mistake is written beside it in the key, in the words you would use to a colleague. Where the mistakes come from, in order: what this teacher has watched her own class do, which is why it is question five above; then the predictable slips — the operation reversed, the step skipped, the unit dropped, the first-listed number used; then the ones documented for that topic at that age. A filler option — the joke, the silly one, "none of the above" — leaves an item with three options working and one spending the child's reading time on nothing. Mechanics, so that the item measures the content rather than test-craft: - **One option is right, and right beyond argument.** If a careful child can build a case for another, the item has two answers and punishes the careful child. - **One idea per item.** An item asking two things produces a wrong answer you cannot attribute. - **Options built to the same grammar and roughly the same length.** The longest one is the answer often enough that children learn to pick it, and then the item measures that. - **The stem stands alone.** Cover the options and read it. If there is no question there, there is no item. - **Avoid NOT and EXCEPT.** Where one is unavoidable, set it in bold. In grades 1–3, do not use them at all; a negation costs more working memory than the content being assessed. - **No item gives away another.** Item 6's stem routinely answers item 2. ## Is it measuring the thing? The check that matters most The classic failure is the science test that measures reading. It is not rare and it does not look like a mistake: the questions are about photosynthesis, the sentences are long, and the children who score badly are the ones who read slowly. The teacher reads the result as a science result, reteaches photosynthesis, and nothing changes. In grades 1–4 this is the failure, not one of them. A grade 2 subtraction check written in word problems sorts the class by reading, and a child who can regroup perfectly is recorded as a child who cannot. Four passes over the draft, in this order: - **Reading load.** Read every stem as the slowest reader in the class. Where a sentence is longer or denser than what that child reads in that subject, cut it down. In grades 1–4, assume the paper will be read aloud, and check it works when it is. - **Vocabulary from outside the content.** *Altogether, represents, determine, in terms of.* A child who cannot decode the word cannot show what she knows about the thing. Where the word *is* the content it stays and is meant to be hard; where it is scaffolding around the content, swap it. - **Multi-step instructions.** "Read the passage, choose two examples and explain how each one supports the author's argument" is three tasks in one sentence, and a child who does two of them is scored as understanding a third of the content. Number the parts. - **Anything never taught.** Not the content — the skill the response demands. Drawing a labelled diagram, writing a paragraph, reading a table, using a protractor. If the class has not done it, the item measures it. Ask which parts of the response format they have practised. The second failure is quieter: the assessment covers less of the outcome than the outcome claims. Three recall items standing in for an outcome about explaining is the common shape. The mark then means something much narrower than its label — and the label is what reaches the report card. Then say in one sentence what the assessment measures. Where it measures two things and both are wanted — the science and the writing — keep both and mark them separately, so a child who understands it and cannot write it down does not get one number hiding both halves. ## The answer key, always Nothing goes to a teacher without one. Per item: the correct answer, one line on why it is correct, and for each distractor the mistake it catches. For short answers, what earns credit — the ideas that have to be present, in words a nine-year-old might use, plus one or two answers that look wrong at a glance and are right. That last part stops the mark scheme penalising the child who thought about it. Then the part that gets left out: **what to do with the pattern.** A cluster on one wrong option is a lesson; the same children spread evenly across three is a class that guessed, and the item either came too early or the teaching never happened. Write down the two or three patterns worth watching for, and what each would mean. ## Rubrics, for work a mark scheme cannot mark A rubric is what you need where the response is a piece of work rather than an answer: writing, a project, an explanation, a thing built. **Levels describe what is on the page.** *Excellent, good, fair, needs improvement* is a scale of the marker's feeling about the work. It tells a child nothing she can act on, and two markers cannot disagree about it usefully, because there is nothing there to disagree about except taste. Every vague quantity word does the same damage: *some, several, thorough, adequate.* Ask what would be visible, and write that instead. **Adjacent levels differ in kind, not in amount.** "Uses some evidence" and "uses more evidence" is one level counted twice, and a child on the lower one learns only that she should do more of whatever it was. What separates real levels is a different sort of work: the thing is absent, the thing is present but doing nothing, the thing is doing its job, the thing is chosen deliberately from among options. **Criteria have to be able to disagree.** If a child scores the same on every criterion every time, the criteria are one criterion wearing several hats. The test is whether you can picture the work that is high on one and low on another. **As many levels as you can actually tell apart, and not one more.** A fifth level nobody can distinguish makes the rubric less reliable, not more precise. Where the school's reporting scale already has levels (`school/gpa-context.md#report-card-cycle`), use those so rubric and report card speak one language; where blank, ask before choosing a number, and never attach marks to the levels yourself. **Write the child's version too.** Second person, her vocabulary, handed out before she starts rather than returned with the mark — most of a rubric's value. **The two-marker test, before you hand it over.** Could a colleague who was not in the room mark the same piece of work and land on the same level? Where the answer is no, find the decision point — it is nearly always one boundary, not all of them — and describe what a piece of work sitting exactly on it looks like. ## Fold in the earlier work Offer this every time, in one line: three or four items from units already finished, sitting alongside the ones on this week's topic. Why that beats a fourth item on the current topic. This week's material is met again tomorrow, and the day after, and in the homework; it is not going anywhere. The unit from October is being forgotten right now and nothing in the week touches it. An item there does two jobs at once — it tells the teacher what survived, and the act of retrieving it makes it survive longer. That second job is the testing effect, and it is why the item is not merely a measurement. Two conditions on it: - **Closed book, from memory.** A question answered by looking something up is not retrieval and does neither job. - **The old items do not count against a mark unless the children were told the assessment was cumulative.** That is a fairness point rather than a memory one, and it is the one teachers get wrong. Where the mark goes on the record, say so in the instructions, in advance. ## When it is one child rather than the class **Where the teacher wants to understand one struggling child, that is `supporting-students`.** The method there is different — put the child's own wrong answers in front of her and ask what she was thinking — and a better quiz will not get there. Say so and route. Two smaller boundaries: what the unit should assess at all is `planning-curriculum`, and results that have to become something a family reads are `communicating-with-families`. ## The part a child reads: `writing-in-a-human-voice` Most of what this skill makes should read flat. A mark scheme is a checklist and an answer key a lookup table; warming either up only makes it slower to use. Three things here are not that, and they all fail the same way if they read like a machine wrote them: - **Rubric level descriptors, and the child's version above all.** A level a child cannot read is a level she cannot use to close the gap, which was the entire reason for describing levels instead of labelling them. *"Demonstrates a developing ability to utilise textual evidence to support inferences"* is a level no child has ever acted on, and a parent reading it on a marked piece of work learns nothing either. - **Item stems the child reads for herself.** Stiff phrasing in a stem is not a style problem, it is reading load — the validity pass and the voice pass are the same pass on a stem, arriving from two directions. - **Anything that travels.** The note that goes home with the result, the sentence about it in a report card comment. So run `writing-in-a-human-voice` over those three, and not over the key, the mark scheme, or the mapping table. Two limits on what it may change: a level's meaning stays where it was — plainer wording is the job, a softer boundary is not — and its invent-nothing rule lands here as a ban on presenting invented work as some real child's, not on the made-up example a level descriptor needs. ## A worked check — grade 5 science, ten minutes, mid-unit Everything in the scenario is the teacher's. Her outcome, in her words: *explain why the moon looks like a different shape at different times of the month.* Grade 5, ten minutes, on a Wednesday with two lessons of the unit left. Purpose: find out who to reteach — nothing here is marked. She has seen two children say the shapes are the Earth's shadow. **Item 1.** Half of the moon is always lit by the sun. Why do we see different shapes from Earth? | | Option | What choosing it means | | --- | --- | --- | | a | The Earth's shadow falls across part of the moon. | The shadow model. Phases and eclipses are one thing to this child, and it is the most common wrong model there is — documented in adults as well as children. | | b | We see the lit half from a different angle each night. | **Correct.** | | c | Clouds in the way hide part of the moon from us. | Something in between blocks it — the child has a covering model rather than a lighting one. | | d | The moon spins, so a different side faces the sun. | Rotation confused with orbit. Plausible because she has heard that the moon spins, and it is true; it is not what makes the shapes. | **Item 2.** Ravi says the moon is only in the sky at night. Kira says she saw it at two in the afternoon. Who is right, and how could you check? *This item was cut.* It is a good item and it is not this outcome — she can answer it while holding the shadow model intact. It goes in a different check or it goes nowhere. Cutting it is the mapping doing its work. **Item 3.** Six pictures, each showing which side of the moon is lit: new, crescent, half, full, then half and crescent again lit on the other side. Put them in the order the moon goes through, starting from new. Six and not four, because the shapes alone are symmetrical — with four, a child running the cycle backwards produces the same list as a child who has it right, and the item reports a misconception it never tested. With the waning pair in, the error it catches is real: a child who puts the two crescents together and the two halves together is sorting by how much is lit rather than by sequence. **Item 4.** Draw the sun, the Earth and the moon at a moment when someone on Earth sees a full moon. Then write two or three sentences saying why the moon looks fully lit from where she is standing. The only item assessing the outcome as she stated it; 1 and 3 assess the pieces it is made of. Ten of item 1 would look thorough and measure recognition. **Mapping.** | Item | Evidence for | Asks the child to | | --- | --- | --- | | 1 | What changes is our angle, not the moon | Pick the mechanism | | 3 | The order of the phases | Sequence | | 4 | Explain why — the outcome itself | Draw and explain in her own words | **Validity notes, and one of them changes the paper.** *"Half of the moon is always lit by the sun"* is a needed premise, not reading load, and it stays. *Waxing* and *waning* appear nowhere, because this check is about the mechanism and those words would sort the class by vocabulary. Item 4 asks for a diagram — **if this class has only ever looked at diagrams and never drawn one, item 4 measures drawing.** Ask her. If they have not, the diagram gets modelled first, or item 4 becomes spoken: she tells it to the teacher with the three objects on the desk in front of her, which in grade 5 is often the better version anyway. **What to do with the pattern.** Everyone on (b) — the misconception is not loose in this room and the last two lessons can proceed. A cluster on (a) — the shadow model is live, and it does not go away by being told; it goes away by being made to predict something and getting it wrong, so the reteach is the model on the desk with a lamp, not a better explanation. ## A worked rubric — grade 6, explaining someone else's mistake The task: here is another child's work on a subtraction problem, and the answer is wrong. Say what went wrong and fix it. Two criteria, because a child can pick the work up at the right place without ever saying what the mistake was — the independence test above, applied to this rubric. | | Finding the error | Fixing it | | --- | --- | --- | | **1** | Says the answer is wrong, or works the problem out correctly alongside, without pointing at any step of the original. | A corrected answer appears with no working of its own. | | **2** | Points at a step and says it is wrong there, whether or not it is the step the error starts at. | Redoes the problem from the beginning, correctly, without returning to the point where it goes wrong. | | **3** | Points at the step where it first goes wrong and names what was done there: "she took 3 from 8 instead of 8 from 3." | Picks up at the point where it goes wrong and carries on from there, and the corrected answer follows from that change. | | **4** | Names that step, names what was done, and says what the other child seems to have believed: "she thinks you always take the smaller digit off the bigger one." | Does that, then checks the new answer another way — adds it back, or estimates — and says the check agrees. | **Why the levels are tellable apart.** Each is something a marker can point to on the page: nothing there, a location, a location plus a named action, a named action plus the belief behind it. No level says *thorough* or *good*. The step from 3 to 4 is not more explaining but a different object — 3 describes what the hand did, 4 describes what the head was doing, and a child never asked for the second will not produce it by trying harder. **The child's version**, which goes out before she starts: > **Finding it.** 1 — you say it is wrong. 2 — you point to a line and say it > goes wrong there. 3 — you point to the line where it *first* goes wrong and > say what she did there. 4 — you say what she did, and what she probably > thought the rule was. > > **Fixing it.** 1 — you write the right answer. 2 — you do the whole thing > again yourself from the start. 3 — you pick it up at the line where it goes > wrong and carry on from there. 4 — you do that, then check it another way and > show that the check agrees. **Calibration — the boundary two markers will split on** is 2 against 3 on the first criterion, every time. The question that settles it: does she name the *action*, or only the *place*? "It goes wrong in the ones column" is 2. "In the ones column she took the smaller digit off the bigger one" is 3. Underline the verb in what the child wrote; if there is not one, it is a 2. A named action at the wrong step is also a 2 — 3 needs both. **What is deliberately not here:** what these levels are worth, what a 3 counts as, and whether the two criteria are averaged. That is the school's, at `school/gpa-context.md#report-card-cycle`. Ask before attaching a number. ## How to work 1. Ask the purpose first, then the five things, in one message. 2. Read the three anchors. Ask for what they leave blank, and leave marks and cut-offs off the instrument rather than inventing them. 3. Write the items to the verb in the outcome, and the distractors to named mistakes. Write the key as you write each item, never afterwards — a key written afterwards is where the ambiguous second right answer survives. 4. Run the four validity passes on the draft: reading load, vocabulary, multi-step instructions, anything never taught. Say what you changed. 5. Build the item-to-outcome mapping and read it for the piece of the outcome nothing assesses. 6. If you can dispatch a subagent, have one answer the items cold without the key and report what it thinks each item is measuring; otherwise do that pass yourself, in sequence, before showing anything. Where its answer is not the outcome, the item is what changes. 7. Run `writing-in-a-human-voice` over the rubric levels, the child's version, and any stem she reads herself. Not over the key. 8. Offer the earlier-units items, in a line, with the reason. If a roster or gradebook tool is available, use it for the class list and the marks; otherwise ask the teacher, and never invent a name, a mark or a date. ## Handing it back If you can write files, put the assessment and its key there and say where they went, because both get opened again next year; otherwise give both in the conversation, ready to copy. Keep the child's copy and the key visibly apart — a key pasted under the items gets photocopied with them. Alongside it, never inside it: - **What it measures**, in one sentence, and what it does not. - **What you assumed**, especially anything read off a blank anchor — above all a marking scale you were not given. - **The one question that would most improve it.** Usually it is what she has seen them get wrong, and it usually rewrites two distractors.