Analyzing assessment data: reliability and item analysis
A nine-question essay on sample test statistics where every answer has to quote a number from a named textbook table — drawn from two different chapters in two different books — and end in something an instructor would actually do.
Editorial process
Last reviewed · August 8, 2026
What does 'analyze the sample test statistics' actually require?
This is a data-interpretation essay wearing the clothes of a concepts essay, and the difference decides the grade. Every question is anchored to supplied numbers — the brief says *use the sample statistics data from the textbooks to respond*, and the first question spells the standard out: is this test reliable, and **what evidence from the statistics supports your answer**. A correct textbook definition of reliability with no number attached has answered half a question. The pattern to repeat nine times is the same: define the concept in a sentence, quote the actual figure from the table, say what that particular figure means for this particular test, and state what an instructor would do about it. Definitions are cheap here; the interpretation is what is being assessed. Write the four moves out as a template once and reuse it, so no answer quietly loses its evidence step.
Note that the questions draw on **two different textbooks**, and mixing them up is a quiet way to lose marks. Questions 1 to 4 use the sample test statistics in Chapter 24 of *Teaching in Nursing: A Guide for Faculty*; questions 5 to 9 use Chapter 11 of *The Nurse Educator's Guide to Assessing Learning Outcomes*. Those are different datasets, so a range or a standard error carried over from the first table into an answer about the second is simply the wrong number. Open both chapters before you start and label your working notes by source. It is also worth doing the arithmetic yourself rather than quoting the textbook's summary sentence, because several of these questions ask what the statistic *provides* rather than what it *is*. Note which chapter each figure came from as you transcribe it, not afterwards from memory.
Standard error of measurement is the question most often answered without content, and it has a trap in its wording. The SEM describes the precision of an *individual student's* score — the band within which their true score probably falls — rather than the quality of the test as a whole, which is what students tend to write. The brief then asks whether this test has a small or large SEM, and small or large is meaningless in the abstract: an SEM of 3 is tight on a 100-point exam and disastrous on a 20-point quiz. Judge it against the score scale and the spread of the raw scores, say what you are comparing it to, and then draw the practical conclusion — most usefully, what it means for students clustered near a pass mark. That last point is usually the most useful sentence in the whole essay for a working instructor.
The item-analysis questions are where you can show genuine judgement, because the statistics are advisory rather than decisive. Difficulty index and discrimination index between them flag items worth examining, and the two interact: research on health-professions examinations finds a negative correlation between average difficulty and average discrimination, so easier items tend to discriminate less well. But thresholds do not settle anything on their own. A six-year analysis of over fourteen thousand items found that standard psychometric cut-offs identified about 85.7% of flawed questions and missed 14.3% — with multiple correct answers and incorrect answer keys the commonest defects. The right answer to *what does an instructor do next* is therefore expert review, prompted by the statistic rather than replaced by it. Naming the two commonest defects also tells the instructor what to look for during that review.
Two practical constraints shape the writing. Nine questions inside 1,000–1,250 words leaves roughly 110 to 140 words each, which is enough for the four-move pattern above and nothing more — so there is no room for an introduction that explains why assessment matters, and none for a conclusion that summarises what you just said. Second, keep the instructor's perspective throughout: three of the questions ask explicitly how an instructor would *use* the information, and that is the difference between describing a statistic and demonstrating you could act on one. Answer in the order asked, signpost each question, and let the numbers carry the argument. Numbering the answers one to nine is the simplest way to make coverage visible, and it costs nine words in total rather than a paragraph of signposting prose.
Question | Source chapter | What a full answer contains |
|---|---|---|
Reliability, and the evidence for it | Ch. 24, Teaching in Nursing | The coefficient itself, plus what that value means for this test |
Trends in the raw scores | Ch. 24 | The pattern, and the instructional decision it supports |
The range, and why it matters | Ch. 24 | The computed range and what spread implies about discrimination |
Standard error of measurement | Ch. 24 | Precision of an individual score, judged against the score scale |
Analysing individual items | Ch. 11, Assessing Learning Outcomes | Difficulty and discrimination as flags for review, not verdicts |
A flawed exam question | Ch. 11 | Review-led remedies — rekey, drop, accept multiple answers |
Likely learning objectives
Inferred from the brief — check these against your own rubric.
- 01Interpret psychometric statistics against a specific dataset rather than defining them abstractly.
- 02Judge a measure of precision relative to the scale it is measured on.
- 03Distinguish a statistical flag from a decision about an assessment item.
- 04Translate measurement findings into instructional action.
Read the full question
Review every instruction before using the planning guidance that follows.
Course-wide instructions that accompany this question
You must proofread your paper. But do not strictly rely on your computer’s spell-checker and grammar-checker; failure to do so indicates a lack of effort on your part and you can expect your grade to suffer accordingly. Papers with numerous misspelled words and grammatical mistakes will be penalized. Read over your paper – in silence and then aloud – before handing it in and make corrections as necessary. Often it is advantageous to have a friend proofread your paper for obvious errors. Handwritten corrections are preferable to uncorrected mistakes. Use a standard 10 to 12 point (10 to 12 characters per inch) typeface. Smaller or compressed type and papers with small margins or single-spacing are hard to read. It is better to let your essay run over the recommended number of pages than to try to compress it into fewer pages. Likewise, large type, large margins, large indentations, triple-spacing, increased leading (space between lines), increased kerning (space between letters), and any other such attempts at “padding” to increase the length of a paper are unacceptable, wasteful of trees, and will not fool your professor. The paper must be neatly formatted, double-spaced with a one-inch margin on the top, bottom, and sides of each page. When submitting hard copy, be sure to use white paper and print out using dark ink. If it is hard to read your essay, it will also be hard to follow your argument
What all nine answers have to contain
- 01Answers to all nine questions, in the order the brief lists them.
- 02Statistics drawn from Chapter 24 of Teaching in Nursing for questions one to four.
- 03Statistics drawn from Chapter 11 of The Nurse Educator's Guide for questions five to nine.
- 04Specific figures quoted as evidence, not only definitions.
- 05An instructor-use implication wherever the brief asks for one.
- 061,000–1,250 words, in the required academic format.
Working through reliability, range, SEM and item analysis
Reliability and the evidence for it
Define reliability, report the coefficient from the sample, and judge this test against it.
Raw score trends and their instructional use
Describe the distribution pattern and what it tells an instructor about the cohort or the test.
Range and standard error of measurement
Report both, and interpret the standard error relative to the score scale.
The process of item analysis
Set out how difficulty and discrimination indices are computed and read together.
Handling a flawed item
Explain the review-led options once a statistic flags a question.
Using two textbook datasets without mixing them
Recommended databases
- The two assigned textbooks, Chapter 24 and Chapter 11
- PubMed Central
- ERIC
- CINAHL
Search sequence
- 1.Open both assigned chapters and transcribe the two sample datasets into separate notes before writing, so no figure is carried across from the wrong table.
- 2.Compute the range and check the reported reliability coefficient yourself, since several questions ask what a statistic provides rather than reciting what it is.
- 3.Read a current item-analysis study to get the relationship between difficulty and discrimination right, because the two are not independent and the essay is stronger for saying so.
- 4.Find evidence on how reliably psychometric thresholds detect flawed items, so the final answer can recommend expert review on grounds rather than as a platitude.
Psychometric evidence for item analysis
These are authoritative starting points, not a ready-made bibliography. A qualified reviewer must confirm that each source fits the assignment and supports the claim beside which it is cited.
Nothing here is cleared for citation until you have read it.
- 01
Item analysis: the impact of distractor efficiency on the difficulty index and discrimination power of multiple-choice items
BMC Medical Education · 2024
Quantifies how the three item statistics interact — distractor efficiency correlates negatively with the difficulty index (r = -0.548) and weakly negatively with discrimination. Use it for the item-analysis questions to show the indices are related rather than independent, and to explain why an item can look acceptable on one measure and not on another.
- 02
Item difficulty index, discrimination index, and reliability of the 26 health professions licensing examinations in 2023, Korea: a psychometric study
Journal of Educational Evaluation for Health Professions · 2024
Real benchmark values from 26 health-professions examinations, including the nursing paper, with reliability at Cronbach's alpha 0.855 or higher across all tests. This gives you an external comparison for judging whether the sample's reliability coefficient is good, and it reports the negative correlation between average difficulty and average discrimination.
- 03
Detection of flawed multiple-choice questions in preclinical medical education using item difficulty and discrimination indices: a six-year analysis
BMC Medical Education · 2025
The strongest evidence for the final questions: across 14,238 items, psychometric thresholds caught about 85.7% of flawed questions and missed 14.3%, with multiple correct answers and incorrect keys the commonest defects. It is what turns 'the instructor should review the item' from an assertion into a supported recommendation.
Before the assessment data essay is submitted
Common mistakes
- Defining reliability correctly and never citing a value from the sample statistics.
- Using one textbook's dataset to answer questions assigned to the other.
- Describing the standard error of measurement as a property of the test rather than of an individual score.
- Calling the standard error small or large without saying what it is being compared to.
- Treating difficulty and discrimination thresholds as verdicts rather than as flags for review.
- Answering 'what trends are seen' without saying what an instructor would do about them.
- Spending scarce words on an introduction and conclusion the question format does not need.
- Merging the nine answers into continuous prose so coverage cannot be checked.
Submission checklist
- All nine questions are answered and individually identifiable.
- Each answer quotes at least one figure from the correct textbook table.
- Chapter 24 data is used for questions one to four and Chapter 11 data for five to nine.
- The reliability answer states the coefficient and interprets its value.
- The standard error answer is framed around individual score precision and a stated comparison.
- Item analysis names both difficulty and discrimination and treats them as advisory.
- Every 'how would an instructor use this' prompt has an explicit answer.
- Word count falls between 1,000 and 1,250.
Use this guide to plan and review your own work. Follow your institution's rules and read our academic-integrity policy.

Written by
Aaron Bishop
MA, Education
assignment interpretation and research-methods coaching across disciplines
Aaron leads the EssayCrackers editorial desk. He works on how assignment briefs are read — what a rubric is actually asking for, and where students most often answer a different question than the one set.

Reviewed by
Dr. Nathan Cole
PhD, Rhetoric & Composition
Argumentation and thesis development
Nathan teaches first-year composition and directs a university writing center. He reviews EssayCrackers guides for argumentative soundness and citation accuracy.