Assessment and Evaluation
What to remember
- Measurement gives a number, assessment collects evidence about learning, and evaluation makes a judgement using that evidence. Test < measurement < assessment < evaluation in scope.
- Formative assessment (assessment FOR learning) happens during teaching to improve it; summative assessment (assessment OF learning) happens at the end to grade and report.
- Continuous and Comprehensive Evaluation (CCE) covers scholastic and co-scholastic areas through regular, multiple tools. A good tool is valid, reliable, objective and usable.
Basic terms
| Term | Meaning | Example |
|---|---|---|
| Test | A set of questions to obtain a response | A 20-mark unit test |
| Measurement | Giving a number to a quality | 15 out of 20 |
| Assessment | Collecting information about learning in many ways | Test, project, observation |
| Evaluation | Judging worth or quality against a standard | "Needs remedial help" |
Purposes of assessment: find what the learner knows, give feedback, plan teaching, identify difficulties, motivate, report progress, and certify. Under the RTE Act 2009, Section 29 asks for continuous and comprehensive evaluation of the child's understanding and ability to apply knowledge, and the child's all-round development. The Act also bars board examinations in elementary classes. It originally barred detention up to Class VIII; the 2019 amendment lets the government allow holding back in Classes V and VIII after a re-examination. The NEP 2020 moves toward competency-based, low-stakes assessment, with a holistic progress card showing self, peer and teacher assessment.
Types of assessment
By purpose and time.
- Placement: before teaching, to decide where to start (entry behaviour).
- Diagnostic: finds the cause of repeated learning difficulties.
- Formative: ongoing; gives feedback; not for ranking.
- Summative: end of unit, term or year; gives grades or marks.
By the standard used.
- Norm-referenced: compares a learner with other learners (rank, percentile).
- Criterion-referenced: compares with a fixed standard or competency (mastery).
- Self-referenced: compares with the learner's own earlier performance.
By who assesses: teacher assessment, peer assessment, self assessment.
| Formative | Summative |
|---|---|
| During learning | After learning |
| Improves teaching and learning | Grades and certifies |
| Frequent, small, informal | Occasional, large, formal |
| Feedback is the key output | Marks or grades are the key output |
| Low stakes | High stakes |
Tools and techniques
Tests: written (essay, short answer, objective), oral, practical. Objective-type items: multiple choice, true-false, matching, fill in the blank. Essay questions test organisation and expression but are subjective in scoring.
Non-test tools:
- Observation schedule and checklist: Yes/No list of behaviours or skills.
- Rating scale: degree of a trait on a scale (for example 1 to 5).
- Rubric: a scoring guide with criteria and performance levels; makes marking transparent. Analytic rubric scores each criterion; holistic rubric gives one overall score.
- Anecdotal record: short factual note about a significant incident.
- Portfolio: purposeful collection of the learner's work over time; shows growth.
- Project, assignment, field work, experiment, seminar, quiz, role play.
- Cumulative record card: long-term record of progress and personal data.
- Sociometry: maps friendship choices in a class (sociogram).
- Questionnaire, interview, inventory: for interests, attitudes and personality.
Question paper planning. A blueprint (table of specification) fixes weightage to objectives (knowledge, understanding, application, skill), content units and question types, so that the paper is balanced.
Bloom's taxonomy (revised): Remember, Understand, Apply, Analyse, Evaluate, Create. Questions at each level use different verbs: define, explain, solve, compare, justify, design.
Qualities of a good tool
| Quality | Meaning |
|---|---|
| Validity | Measures what it claims to measure (content, criterion, construct, face) |
| Reliability | Gives consistent results (test-retest, split-half, parallel forms, KR-20) |
| Objectivity | Different scorers give the same score |
| Usability / practicability | Easy to administer, score and interpret; affordable |
| Discrimination | Separates good from weak learners |
| Comprehensiveness | Covers the whole content and objectives |
A valid test is always reliable, but a reliable test need not be valid.
Continuous and Comprehensive Evaluation (CCE)
- Continuous: assessment is part of teaching, regular and spread over the year, using many occasions and tools.
- Comprehensive: covers scholastic areas (subjects) and co-scholastic areas (life skills, attitudes and values, sports, arts, work experience, health). It uses varied tools.
- FA and SA: In the CCE pattern used in AP schools, formative assessments (FA) are done through the year as written work, projects, assignments, slip tests, and so on, while summative assessments (SA) are term-end written papers. The number of FAs and SAs and their weightage have changed over the years; check the latest official orders.
- Grading: a grade scale (for example A to E) replaces raw marks to reduce stress and unhealthy competition.
- Records: teachers maintain FA records, a cumulative record and a progress card shared with parents.
- Benefits: reduces exam fear, finds problems early, assesses all-round growth, and links teaching with testing.
- Challenges: teacher workload, large classes, subjective scoring and record keeping.
Analysis of learner data
Descriptive statistics.
- Mean = sum of scores ÷ number of scores.
- Median = middle value after arranging in order (average of two middle values if even).
- Mode = most frequent score.
- Range = highest − lowest.
- Standard deviation measures spread around the mean; a small SD means scores are close together.
- Percentile rank: percentage of learners scoring below a given score.
*Worked example.* Scores 6, 8, 8, 10, 12: sum = 44, mean = 44 ÷ 5 = 8.8; median = 8; mode = 8; range = 12 − 6 = 6.
Item analysis.
- Difficulty index (P) = (number answering correctly ÷ total number) × 100. Values near 50 are ideal; very high means easy, very low means hard.
- Discrimination index (D) = (correct in upper group − correct in lower group) ÷ number in one group. The upper and lower groups are usually the top and bottom 27 percent. D ranges from −1 to +1; positive and high is good.
- Distractor analysis checks whether wrong options attract weak learners.
*Worked example.* In upper and lower groups of 10 each, 8 and 4 learners got an item right. D = (8 − 4) ÷ 10 = 0.4. Difficulty = (8 + 4) ÷ 20 × 100 = 60 percent.
Using the data. Error analysis shows common mistakes. Teachers give remedial teaching to learners below the expected level and enrichment to fast learners. Results are shared as constructive feedback: specific, timely, and focused on how to improve, not on blame. Data from class results helps plan re-teaching and revise the teacher's own methods. Frequency tables, bar graphs, histograms and line graphs show data clearly.
Feedback, grading and classroom practice
Feedback is information given to a learner about performance. Good feedback is timely, specific, kind and tells the learner what to do next. Written comments such as "Your steps are correct; check the units" help more than a mark alone. Feedback can be given by the teacher, by peers, or by the learner through self checking.
Grading. Marks are converted to grades by fixed ranges. In direct grading the teacher awards a grade straight from performance; in indirect grading marks are first given, then changed to grades. Grades remove small differences, such as one mark, and reduce pressure.
Merits and limits of question types.
| Type | Merit | Limit |
|---|---|---|
| Essay | Tests expression, organisation, higher thinking | Subjective scoring, low coverage |
| Short answer | Fair coverage, quick to answer | Limited depth |
| Objective | Fast, objective scoring, wide coverage | Guessing, hard to write well |
| Practical / oral | Tests skills and communication | Time consuming |
Good classroom practice. Share the learning outcomes and criteria with learners before the task. Use open questions and wait time. Let learners assess their own and a friend's work using a simple rubric. Avoid comparing a child with others in public. Keep a record of each child, including children with special needs, and use suitable adaptations such as extra time, large print or oral answers.
Exam traps
- Assessment is wider than measurement; evaluation is the judgement stage.
- Formative = FOR learning; summative = OF learning.
- Norm-referenced compares with others; criterion-referenced compares with a standard.
- Diagnostic assessment finds causes of difficulty; placement assessment finds starting point.
- Reliability does not guarantee validity.
- Rubric is a scoring guide; portfolio is a collection of work; rating scale shows degree; checklist shows presence or absence.
- Co-scholastic areas are not "extra marks"; they are part of comprehensive evaluation.
- High mean does not mean small SD; SD shows spread, not the average.
One-liners
- 1. Evaluation = measurement + value judgement.
- 2. RTE Section 29 mentions continuous and comprehensive evaluation.
- 3. Formative assessment is low-stakes and gives feedback.
- 4. Summative assessment comes at the end of a term or year.
- 5. A blueprint keeps a question paper balanced.
- 6. Sociometry studies social relations in a group.
- 7. An anecdotal record is a short note of an incident.
- 8. A portfolio shows growth over time.
- 9. Mode is the most frequent score.
- 10. Difficulty index near 50 percent is ideal.
- 11. Discrimination index ranges from −1 to +1.
- 12. Remedial teaching helps learners who are behind.
Practice questions
Which of these gives only a number to a learner's performance?
- Feedback
- Diagnosis
- Evaluation
- Measurement
Answer
D. Measurement
Measurement is quantitative; evaluation adds judgement.
Assessment that takes place during teaching to improve learning is
- formative assessment
- summative assessment
- placement test only
- board examination
Answer
A. formative assessment
Formative assessment gives feedback while learning continues.
Assessment at the end of a term to give grades is
- formative assessment
- self assessment
- diagnostic assessment
- summative assessment
Answer
D. summative assessment
Summative assessment sums up achievement.
Comparing a learner's score with those of other learners is
- self-referenced assessment
- criterion-referenced assessment
- diagnostic assessment
- norm-referenced assessment
Answer
D. norm-referenced assessment
Norm referencing ranks learners against a group.
Comparing a learner's performance with a fixed standard of mastery is
- sociometry
- norm-referenced assessment
- criterion-referenced assessment
- percentile ranking
Answer
C. criterion-referenced assessment
Criterion referencing uses a set standard.
Assessment used to find the cause of repeated learning difficulty is
- diagnostic assessment
- norm-referenced assessment
- placement assessment
- summative assessment
Answer
A. diagnostic assessment
Diagnostic tests locate the root of the problem.
Which RTE Act section refers to continuous and comprehensive evaluation?
- Section 21
- Section 12
- Section 16
- Section 29
Answer
D. Section 29
Section 29 deals with curriculum and evaluation procedure.
A scoring guide listing criteria and levels of performance is a
- rubric
- sociogram
- cumulative record
- checklist
Answer
A. rubric
A rubric makes scoring transparent.
A purposeful collection of a learner's work over time is a
- anecdotal record
- rating scale
- portfolio
- blueprint
Answer
C. portfolio
Portfolios show growth.
A short factual note on a significant incident about a child is a(n)
- blueprint
- anecdotal record
- rubric
- question bank
Answer
B. anecdotal record
Anecdotal records note actual incidents.
A sociogram is drawn using
- sociometry
- blueprint
- item analysis
- rubric
Answer
A. sociometry
Sociometry maps choices and relations in a group.
A table that fixes weightage to objectives, content and question types is a
- sociogram
- rating scale
- portfolio
- blueprint
Answer
D. blueprint
A blueprint ensures a balanced question paper.
The test quality of measuring what it claims to measure is
- usability
- objectivity
- validity
- reliability
Answer
C. validity
Validity is about the purpose of the test.
Consistency of test results is called
- discrimination
- reliability
- validity
- difficulty
Answer
B. reliability
Reliable tests give stable scores.
Test-retest and split-half are methods of finding
- reliability
- validity
- grade
- difficulty index
Answer
A. reliability
These are standard reliability estimates.
Which of these is a co-scholastic area?
- Mathematics marks
- Written paper
- Unit test
- Life skills
Answer
D. Life skills
Life skills, arts, sports and values are co-scholastic.
Which of these is the highest level in the revised Bloom's taxonomy?
- Analyse
- Remember
- Create
- Apply
Answer
C. Create
Revised order ends with Create.
The 'checklist' tool records
- a full essay of the learner
- presence or absence of a behaviour
- friendship choices
- degree of a trait on a scale
Answer
B. presence or absence of a behaviour
A checklist is Yes/No; a rating scale shows degree.
Extra practice and challenge tasks for fast learners are called
- detention
- remediation
- enrichment
- screening
Answer
C. enrichment
Enrichment extends learning for advanced learners.
Teaching planned for learners who are below the expected level is
- enrichment teaching
- remedial teaching
- team teaching
- drill only
Answer
B. remedial teaching
Remedial teaching closes learning gaps.
The mean of 4, 6, 8, 10 and 12 is
- 8
- 40
- 6
- 10
Answer
A. 8
Sum 40 divided by 5 is 8.
The median of 3, 7, 5, 9 and 11 is
- 9
- 5
- 35
- 7
Answer
D. 7
Sorted 3, 5, 7, 9, 11; the middle value is 7.
The mode of the scores 5, 6, 6, 7, 8, 6, 9 is
- 6
- 5
- 8
- 7
Answer
A. 6
6 occurs three times.
The highest score is 92 and the lowest is 35. The range is
- 67
- 127
- 57
- 46
Answer
C. 57
92 minus 35 equals 57.
In an item, 30 of 40 learners answered correctly. The difficulty index is
- 25 percent
- 75 percent
- 30 percent
- 133 percent
Answer
B. 75 percent
30 ÷ 40 × 100 = 75.
In groups of 10, 9 learners in the upper group and 3 in the lower group got an item right. The discrimination index is
- 0.3
- 0.9
- 1.2
- 0.6
Answer
D. 0.6
(9 − 3) ÷ 10 = 0.6.
The median of the scores 2, 4, 6 and 8 is
- 4
- 6
- 20
- 5
Answer
D. 5
Average of 4 and 6 is 5.
The mean of a class is 40. Each learner's score is increased by 5. The new mean is
- 35
- 45
- 40
- 200
Answer
B. 45
Adding a constant adds the same to the mean.
In upper and lower groups of 15 each, 12 and 6 learners answered an item correctly. The difficulty index is
- 80 percent
- 40 percent
- 60 percent
- 20 percent
Answer
C. 60 percent
(12 + 6) ÷ 30 × 100 = 60.
A teacher shares a rubric with learners before a project so that they know the criteria. This mainly improves
- transparency and self assessment
- only ranking
- only punishment
- only attendance
Answer
A. transparency and self assessment
Known criteria help learners judge their own work.
A teacher writes 'Your steps are right; check the units' on a student's notebook. This is
- descriptive feedback
- a grade
- a norm
- a rank
Answer
A. descriptive feedback
Specific comments tell the learner what to improve.
Most learners got a particular item wrong, and all of them chose the same wrong option. The teacher should first
- punish the class
- analyse the error and re-teach the idea
- lower all marks
- ignore the item
Answer
B. analyse the error and re-teach the idea
Error analysis guides re-teaching.
Consider the statements: 1. Formative assessment is mainly for giving feedback. 2. Summative assessment is done only during teaching.
- 1 only
- 2 only
- Both 1 and 2
- Neither 1 nor 2
Answer
A. 1 only
Summative assessment comes at the end.
Consider the statements: 1. A reliable test is always valid. 2. A valid test is also reliable.
- 1 only
- 2 only
- Both 1 and 2
- Neither 1 nor 2
Answer
B. 2 only
A reliable test may consistently measure the wrong thing.
Consider the statements about CCE: 1. It includes both scholastic and co-scholastic areas. 2. It is based only on one year-end examination.
- 1 only
- 2 only
- Both 1 and 2
- Neither 1 nor 2
Answer
A. 1 only
CCE is continuous and comprehensive.
Consider the statements about the standard deviation: 1. It measures the spread of scores around the mean. 2. A small value means scores are close together.
- 1 only
- 2 only
- Both 1 and 2
- Neither 1 nor 2
Answer
C. Both 1 and 2
Both are correct.
Consider the statements about item analysis: 1. A difficulty index near 50 percent is considered ideal. 2. Discrimination index can range from −1 to +1.
- 1 only
- 2 only
- Both 1 and 2
- Neither 1 nor 2
Answer
C. Both 1 and 2
Both statements are correct.
Consider the statements: 1. Range is the average of the highest and lowest scores. 2. Mode is the most frequent score.
- 1 only
- 2 only
- Both 1 and 2
- Neither 1 nor 2
Answer
B. 2 only
Range is highest minus lowest.
Match the tool with its use: 1. Rubric 2. Rating scale 3. Checklist 4. Portfolio P. Collection of work Q. Presence or absence R. Criteria and levels S. Degree of a trait
- 1-R, 2-Q, 3-S, 4-P
- 1-S, 2-R, 3-P, 4-Q
- 1-P, 2-S, 3-Q, 4-R
- 1-R, 2-S, 3-Q, 4-P
Answer
D. 1-R, 2-S, 3-Q, 4-P
Rubric criteria; scale degree; checklist yes/no; portfolio work collection.
Match the test quality with its meaning: 1. Validity 2. Reliability 3. Objectivity 4. Usability P. Same score by different scorers Q. Consistency R. Measures what it should S. Ease of use
- 1-R, 2-P, 3-Q, 4-S
- 1-Q, 2-R, 3-S, 4-P
- 1-S, 2-Q, 3-P, 4-R
- 1-R, 2-Q, 3-P, 4-S
Answer
D. 1-R, 2-Q, 3-P, 4-S
Standard definitions of the four qualities.
Match the assessment with its purpose: 1. Placement 2. Diagnostic 3. Formative 4. Summative P. Grades at end Q. Starting point R. Cause of difficulty S. Feedback during learning
- 1-Q, 2-S, 3-R, 4-P
- 1-S, 2-R, 3-Q, 4-P
- 1-Q, 2-R, 3-S, 4-P
- 1-R, 2-Q, 3-P, 4-S
Answer
C. 1-Q, 2-R, 3-S, 4-P
Placement finds the start; diagnostic the cause; formative feedback; summative grades.
Consider the statements about the NEP 2020 view of assessment: 1. It supports competency-based, low-stakes assessment. 2. It supports a holistic progress card with self and peer assessment.
- 1 only
- 2 only
- Both 1 and 2
- Neither 1 nor 2
Answer
C. Both 1 and 2
Both reflect the NEP 2020 direction.
Consider the statements about feedback: 1. It should only compare the child with others in public. 2. It should be timely and specific.
- 1 only
- 2 only
- Both 1 and 2
- Neither 1 nor 2
Answer
B. 2 only
Public comparison harms learners.
Consider the statements about grading: 1. Grades reduce stress caused by small mark differences. 2. Grades replace the need to give any feedback.
- 1 only
- 2 only
- Both 1 and 2
- Neither 1 nor 2
Answer
A. 1 only
Feedback remains essential.
Consider the statements about the discrimination index: 1. A high positive value means the item separates strong and weak learners. 2. A negative value means more weak learners than strong learners answered correctly.
- 1 only
- 2 only
- Both 1 and 2
- Neither 1 nor 2
Answer
C. Both 1 and 2
Both are correct.