←
Teaching Mathematics: Methodology · Chapter 8

Evaluation in Mathematics: Tools, Test Characteristics, CCE and SAT

What to remember

  • Evaluation is wider than measurement: it collects evidence (measurement), judges it against objectives and uses the result to improve learning and teaching.
  • A good test is valid, reliable, objective and usable; validity is the most important quality.
  • CCE (Continuous and Comprehensive Evaluation) assesses both scholastic and co-scholastic growth, regularly, through formative assessment and periodic summative assessment (SAT).

Measurement, assessment and evaluation

TermMeaningNature
MeasurementGiving a number to a trait (a score of 18 out of 25)Quantitative only
AssessmentCollecting information about learning through many toolsQuantitative and qualitative
EvaluationJudging the worth of the information against objectives and deciding what to doJudgement plus action
TestA tool or set of questions used to measureInstrument

Evaluation = Measurement (quantitative) + Qualitative description + Value judgement. It is a continuous, cooperative and objective-based process. Its purposes are: knowing the child's progress, finding difficulties, giving feedback, motivating learners, improving teaching methods, grading and reporting, and checking whether the objectives were achieved.

Objectives and the base for evaluation

Evaluation must be linked to objectives. The commonly taught objectives of mathematics are knowledge, understanding, application and skill, and, in the revised Bloom scheme, remember, understand, apply, analyse, evaluate and create. The state curriculum places the five process standards at the centre: problem solving, reasoning and proof, communication, connections, and visualisation and representation. Each question should be written to test one such objective. A blueprint (a table of weightage to objectives, content units and question types) makes the test balanced.

Types of evaluation

TypeWhenPurpose
PlacementBefore teachingFind entry level of the learner
DiagnosticDuring teachingFind the cause of repeated difficulty
FormativeDuring the learning processImprove learning; feedback; "assessment for learning"
SummativeAt the end of a term or yearGrading, promotion; "assessment of learning"

Formative assessment is low-stakes, frequent and informal as well as formal; summative is a formal test of what was learnt. Assessment "as" learning means the child assesses and reflects on his own learning (self- and peer-assessment).

Tools and techniques of evaluation

Tests

  • Written tests: essay type, short answer type, very short answer and objective type (multiple choice, fill in the blanks, true-false, matching).
  • Oral tests: quiz, viva, mental maths.
  • Practical tests: use of laboratory activities, construction and measurement.
  • Diagnostic and achievement tests.

Non-test tools

  • Observation: teacher watches how the child works.
  • Checklist: list of behaviours marked as present or absent.
  • Rating scale: judges the degree of a trait (for example 1 to 5).
  • Anecdotal record: short factual note on an incident.
  • Cumulative record card: long-term record of the child.
  • Portfolio: purposeful collection of a child's work over time.
  • Project, assignment and mathematics laboratory activity.
  • Rubric: scoring guide describing levels of performance for each criterion.
  • Interview, questionnaire, sociometry for personal-social traits.

Types of questions: merits and limits

TypeMeritLimit
Essay or long answerTests organisation, reasoning, expressionLow objectivity, covers little content, time-taking to mark
Short answerWider coverage, fasterLimited depth
Objective typeHighly objective, wide coverage, quick scoringGuessing possible, hard to prepare, tests mainly recall

Characteristics of a good test

  • 1. Validity: the test measures what it is meant to measure. Types: content validity, construct validity, criterion-related (concurrent and predictive) validity, and face validity. A valid test is always reliable, but a reliable test need not be valid.
  • 2. Reliability: consistency of results. Methods: test-retest, equivalent (parallel) forms, split-half, and Kuder-Richardson formulae. Reliability is increased by more items, clear instructions and objective scoring.
  • 3. Objectivity: result does not depend on who marks it. Objective-type tests have high objectivity.
  • 4. Usability (practicability): easy to administer, score and interpret, with reasonable cost and time.
  • 5. Discrimination: ability to separate good and poor students.
  • 6. Difficulty level: items neither too hard nor too easy.
  • 7. Comprehensiveness: covers all units and objectives, as planned in the blueprint.
  • 8. Norms: standards for comparing scores.

Item analysis

  • Difficulty index = (number of students answering correctly ÷ total students tested) × 100. Example: 30 of 40 answer correctly: 30/40 × 100 = 75, an easy item.
  • Discrimination index = (R_U − R_L) ÷ N, where R_U and R_L are the correct answers in the upper and lower groups (commonly top and bottom 27%) and N is the number in one group. Values near +1 discriminate strongly; a negative value shows a faulty item.
  • For a multiple-choice item, distractors should be attractive to weak students.

Steps in preparing a test

Plan the test and objectives, prepare the blueprint, write the items, review and edit them, arrange items from easy to difficult, write clear instructions and a scoring key and marking scheme, try out the test, analyse the items, and finalise.

Diagnostic testing and remedial work

A diagnostic test locates the exact point of weakness (for example regrouping in subtraction) rather than giving only a score. It has many short items on one topic. After diagnosis the teacher provides remedial teaching and re-tests. Error analysis, finding the pattern of errors, is a key technique.

CCE: Continuous and Comprehensive Evaluation

Meaning: "Continuous" means evaluation is a regular, built-in part of teaching, using frequent small assessments with feedback. "Comprehensive" means it covers all aspects of growth: scholastic (subject learning) and co-scholastic (life skills, attitudes, values, art, sports, work education), using many tools.

Basis: The Right to Education Act, 2009, Section 29, requires a comprehensive and continuous evaluation system, and the National Curriculum Framework 2005 recommended reducing examination stress and moving away from rote testing. The Act originally said no child is held back or expelled in elementary education; the RTE (Amendment) Act, 2019 lets states hold back a child in Classes 5 and 8 who fails a re-examination, and the Centre ended no-detention in its own schools in December 2024. Expulsion is still barred.

Features: reduces fear of examinations; regular diagnosis and remediation; uses varied tools; tests understanding and application, not only memory; involves self- and peer assessment; focuses on the whole child; helps the teacher improve teaching.

Structure in the state's classes

  • Formative assessments (FA): several in a year, based on tools such as written work, slip tests, projects, assignments, oral work, observation and portfolio; marks are kept as records through the year.
  • Summative assessments (SA): formal periodic tests, usually mid-year and year-end, based on written papers.
  • The records and grades show the pupil's progress; the focus is on feedback rather than only on marks.

SAT: summative assessment tests

A summative assessment test is the formal written test at the end of a term or year. It checks overall achievement of the syllabus and is used for grading. A blueprint decides weightage of content units, objectives and question types. Its questions should reflect the process standards, include application and reasoning questions, and avoid pure rote recall. Result analysis should be used to plan remedial teaching for the next term.

Comparison table: formative and summative

PointFormativeSummative
PurposeImprove learningGrade or promote
TimeDuring teachingEnd of term or year
MarksLow-stake, recordedHigher-stake
ToolsMany (projects, oral, observation)Mostly written test
FeedbackImmediateLate

Using evaluation results in the classroom

Evaluation is useful only when the results are used. After a test the teacher should discuss common errors with the class, give individual feedback, plan remedial teaching for weak learners and enrichment for fast learners, and revise his own method if many children failed the same item. Feedback should be specific and timely, and comments should say what to improve rather than only give a mark. Records of all assessments should be kept so that growth can be seen over time, and parents should be informed about the child's progress in simple language.

Exam traps

  • Measurement is only quantitative; evaluation includes judgement and action.
  • Validity and reliability are different: a test can be reliable without being valid.
  • Validity is the most important characteristic; objectivity is about scorer-independence.
  • Formative = assessment for learning; summative = assessment of learning.
  • Diagnostic test finds causes; achievement test gives overall level.
  • Checklist records present or absent; rating scale records degree.
  • Portfolio is a collection of work; anecdotal record is a note of an incident.
  • Continuous = regular; comprehensive = all-round (scholastic and co-scholastic).

One-liners

  • 1. Evaluation is a continuous, objective-based process.
  • 2. The most important quality of a test is validity.
  • 3. Test-retest, split-half and parallel forms are methods of finding reliability.
  • 4. Objective-type tests give the highest objectivity.
  • 5. A blueprint gives weightage to objectives, content and question types.
  • 6. Difficulty index = (students correct ÷ total) × 100.
  • 7. Discrimination index = (R_U − R_L) ÷ N.
  • 8. Rubric is a scoring guide with performance levels.
  • 9. RTE Act 2009 Section 29 refers to continuous and comprehensive evaluation.
  • 10. CCE covers scholastic and co-scholastic areas.
  • 11. Formative assessment gives feedback during learning.
  • 12. Diagnostic tests are followed by remedial teaching.

Practice questions

  1. Assigning a number to a trait, such as a score of 18 out of 25, is called

    1. evaluation
    2. curriculum
    3. promotion
    4. measurement
    Answer

    D. measurement

    Measurement is only quantitative; evaluation adds judgement.

  2. Evaluation differs from measurement because evaluation

    1. gives only a number
    2. is done only at the end of the year
    3. uses only written tests
    4. includes value judgement and action to improve learning
    Answer

    D. includes value judgement and action to improve learning

    Evaluation = measurement + qualitative description + judgement.

  3. The most important characteristic of a good test is

    1. length
    2. difficulty only
    3. attractive printing
    4. validity
    Answer

    D. validity

    A test must measure what it claims to measure.

  4. A test that measures what it is intended to measure is said to be

    1. objective
    2. random
    3. valid
    4. usable
    Answer

    C. valid

    This is the definition of validity.

  5. Consistency of test results on repeated use is called

    1. reliability
    2. norm
    3. validity
    4. usability
    Answer

    A. reliability

    Reliable tests give stable results.

  6. Which statement is correct?

    1. A reliable test is always valid
    2. A valid test is never reliable
    3. A valid test is always reliable, but a reliable test need not be valid
    4. Validity and reliability are the same
    Answer

    C. A valid test is always reliable, but a reliable test need not be valid

    Validity needs reliability, but reliability alone does not ensure validity.

  7. Objectivity of a test means

    1. the score does not depend on who marks it
    2. the test has many essay questions
    3. the test is printed neatly
    4. the test is easy
    Answer

    A. the score does not depend on who marks it

    Objective scoring gives the same result for any examiner.

  8. A table that gives weightage to objectives, content units and question types is a

    1. blueprint
    2. checklist
    3. portfolio
    4. rubric
    Answer

    A. blueprint

    A blueprint plans a balanced test.

  9. Evaluation done at the end of a term to grade students is

    1. formative evaluation
    2. diagnostic evaluation
    3. summative evaluation
    4. placement evaluation
    Answer

    C. summative evaluation

    Summative assessment is of learning.

  10. Assessment during the learning process to give feedback and improve learning is

    1. formative assessment
    2. placement test
    3. annual examination
    4. summative assessment
    Answer

    A. formative assessment

    Formative assessment is assessment for learning.

  11. A test used to find the cause of a child's repeated difficulty in a topic is

    1. aptitude test
    2. achievement test
    3. diagnostic test
    4. summative test
    Answer

    C. diagnostic test

    Diagnostic tests locate specific weaknesses.

  12. A list of behaviours in which the teacher marks whether each is present or absent is a

    1. portfolio
    2. rating scale
    3. blueprint
    4. checklist
    Answer

    D. checklist

    Checklists record presence or absence.

  13. A tool that judges the degree of a trait on a scale such as 1 to 5 is a

    1. anecdotal record
    2. diagnostic test
    3. checklist
    4. rating scale
    Answer

    D. rating scale

    Rating scales show levels or degrees.

  14. A short factual note of an important incident in a child's behaviour is called

    1. an anecdotal record
    2. a portfolio
    3. a rubric
    4. a blueprint
    Answer

    A. an anecdotal record

    Anecdotal records capture incidents.

  15. A purposeful collection of a student's work over a period is a

    1. distractor
    2. blueprint
    3. portfolio
    4. checklist
    Answer

    C. portfolio

    Portfolios show growth over time.

  16. A scoring guide that describes levels of performance for each criterion is a

    1. blueprint
    2. rubric
    3. distractor
    4. norm
    Answer

    B. rubric

    Rubrics make scoring of projects and tasks clear.

  17. In a multiple-choice item, the wrong options are called

    1. distractors
    2. stems
    3. rubrics
    4. keys
    Answer

    A. distractors

    The correct option is the key; the others are distractors.

  18. If 30 out of 40 students answer an item correctly, its difficulty index (percentage) is

    1. 40
    2. 30
    3. 75
    4. 25
    Answer

    C. 75

    30/40 × 100 = 75; an easy item.

  19. If 20 of 50 students answer an item correctly, its difficulty index is

    1. 50
    2. 20
    3. 60
    4. 40
    Answer

    D. 40

    20/50 × 100 = 40.

  20. In item analysis, the discrimination index equals

    1. (R_U − R_L) ÷ N
    2. (R_U + R_L) ÷ N
    3. R_U × R_L
    4. N ÷ (R_U − R_L)
    Answer

    A. (R_U − R_L) ÷ N

    Correct answers of the upper group minus the lower group, divided by group size.

  21. A negative discrimination index shows that the item

    1. is faulty because weak students did better than strong students
    2. is perfect
    3. is very easy
    4. is unusable only for oral tests
    Answer

    A. is faulty because weak students did better than strong students

    Good students should do better than poor students.

  22. The Right to Education Act, 2009, which provides for comprehensive and continuous evaluation, does so in Section

    1. 12
    2. 16
    3. 17
    4. 29
    Answer

    D. 29

    Section 29 deals with curriculum and evaluation procedure.

  23. In CCE the word 'comprehensive' refers to

    1. only yearly tests
    2. only oral tests
    3. assessment of scholastic and co-scholastic areas
    4. only Mathematics
    Answer

    C. assessment of scholastic and co-scholastic areas

    Comprehensive means all-round growth, including life skills and attitudes.

  24. In CCE the word 'continuous' refers to

    1. one test at the end of the year
    2. tests only before holidays
    3. a single final examination
    4. regular assessment built into teaching
    Answer

    D. regular assessment built into teaching

    Continuous assessment is frequent, with feedback.

  25. A main purpose of formative assessment in CCE is

    1. to label students as failures
    2. to give feedback so learning can improve
    3. to avoid all records
    4. to detain students
    Answer

    B. to give feedback so learning can improve

    Formative assessment improves learning and teaching.

  26. A summative assessment test (SAT) is

    1. a daily oral question
    2. a self-assessment note
    3. a formal test at the end of a term or year for overall achievement
    4. an observation checklist
    Answer

    C. a formal test at the end of a term or year for overall achievement

    SAT checks overall achievement and is used for grading.

  27. Which tool is best for assessing a mathematics project on collecting and analysing data?

    1. Placement test
    2. Rubric
    3. Single multiple-choice item
    4. Anecdotal record
    Answer

    B. Rubric

    A rubric scores the different criteria of a project.

  28. Assessment 'as' learning refers to

    1. students assessing and reflecting on their own learning
    2. a placement test
    3. teacher punishment
    4. a final examination
    Answer

    A. students assessing and reflecting on their own learning

    Self- and peer-assessment are assessment as learning.

  29. Which of the following is NOT one of the five process standards of mathematics?

    1. Problem solving
    2. Reasoning and proof
    3. Communication
    4. Memorisation of formulas
    Answer

    D. Memorisation of formulas

    The standards are problem solving, reasoning and proof, communication, connections, and visualisation and representation.

  30. A test that is easy to administer, score and interpret has good

    1. validity only
    2. usability (practicability)
    3. objectivity only
    4. negative discrimination
    Answer

    B. usability (practicability)

    Usability concerns time, cost and ease of use.

  31. Consider the statements on formative and summative assessment. 1. Formative assessment is assessment for learning. 2. Summative assessment is assessment of learning.

    1. 1 only
    2. 2 only
    3. Both 1 and 2
    4. Neither 1 nor 2
    Answer

    C. Both 1 and 2

    Both statements are correct.

  32. Consider the statements on test characteristics. 1. Objectivity means the test measures what it intends to measure. 2. A test can be reliable without being valid.

    1. 1 only
    2. 2 only
    3. Both 1 and 2
    4. Neither 1 nor 2
    Answer

    B. 2 only

    Reliable but not valid is possible; the meaning given in the other statement is validity, not objectivity.

  33. Consider the statements on CCE. 1. It covers both scholastic and co-scholastic aspects. 2. It depends on a single final written examination.

    1. 1 only
    2. 2 only
    3. Both 1 and 2
    4. Neither 1 nor 2
    Answer

    A. 1 only

    CCE uses many tools over time; statement 2 is wrong.

  34. Consider the statements on objective-type tests. 1. They allow guessing. 2. They allow wide coverage of content and quick scoring.

    1. 1 only
    2. 2 only
    3. Both 1 and 2
    4. Neither 1 nor 2
    Answer

    C. Both 1 and 2

    Both are correct features of objective-type tests.

  35. Consider the statements on a diagnostic test. 1. It locates the specific weak point of the learner. 2. It is followed by remedial teaching.

    1. 1 only
    2. 2 only
    3. Both 1 and 2
    4. Neither 1 nor 2
    Answer

    C. Both 1 and 2

    Diagnosis leads to remedy.

  36. Consider the statements on essay-type questions. 1. They have high objectivity of scoring. 2. They test organisation of ideas and expression.

    1. 1 only
    2. 2 only
    3. Both 1 and 2
    4. Neither 1 nor 2
    Answer

    B. 2 only

    Essay questions have low objectivity; statement 1 is wrong.

  37. Consider the statements on a checklist and a rating scale. 1. A checklist shows presence or absence of a behaviour. 2. A rating scale shows only present or absent.

    1. 1 only
    2. 2 only
    3. Both 1 and 2
    4. Neither 1 nor 2
    Answer

    A. 1 only

    Rating scales show degree; statement 2 is wrong.

  38. Consider the statements on evaluation. 1. Evaluation is complete when a number is given. 2. Evaluation uses measurement as one of its inputs.

    1. 1 only
    2. 2 only
    3. Both 1 and 2
    4. Neither 1 nor 2
    Answer

    B. 2 only

    Evaluation also needs judgement and action; statement 1 is wrong.

  39. Match the tool with its use. P. Portfolio Q. Rubric R. Checklist S. Anecdotal record 1. Records presence or absence of behaviour 2. Collection of work over time 3. Describes performance levels 4. Notes a particular incident

    1. P-2, Q-1, R-3, S-4
    2. P-1, Q-3, R-2, S-4
    3. P-3, Q-2, R-4, S-1
    4. P-2, Q-3, R-1, S-4
    Answer

    D. P-2, Q-3, R-1, S-4

    Portfolio gathers work; rubric describes levels; checklist records presence; anecdotal record notes incidents.

  40. Match the type of evaluation with its purpose. P. Placement Q. Diagnostic R. Formative S. Summative

    1. P-improve learning, Q-grading, R-entry level, S-find causes
    2. P-grading, Q-improve learning, R-find causes, S-entry level
    3. P-find causes, Q-entry level, R-grading, S-improve learning
    4. P-entry level, Q-find causes of difficulty, R-improve learning, S-grading
    Answer

    D. P-entry level, Q-find causes of difficulty, R-improve learning, S-grading

    These are the standard purposes of the four types of evaluation.

  41. Match the quality of a test with its meaning. P. Validity Q. Reliability R. Objectivity S. Usability

    1. P-scorer independence, Q-ease of use, R-consistency, S-measures what it should
    2. P-consistency, Q-measures what it should, R-ease of use, S-scorer independence
    3. P-ease of use, Q-scorer independence, R-measures what it should, S-consistency
    4. P-measures what it should, Q-consistency, R-scorer independence, S-ease of use
    Answer

    D. P-measures what it should, Q-consistency, R-scorer independence, S-ease of use

    Standard meanings of the four qualities.

  42. A teacher keeps notes on each child's work in class, oral answers and group work over the year. This is

    1. a single summative test
    2. continuous observation-based assessment
    3. a placement test only
    4. an achievement test only
    Answer

    B. continuous observation-based assessment

    Regular observation and records are part of continuous assessment.

  43. A math test has 10 questions all from Chapter 1 although five chapters were taught. This test is weak in

    1. objectivity
    2. usability only
    3. content validity
    4. printing
    Answer

    C. content validity

    It does not cover the whole content planned in the blueprint.

  44. Two examiners give widely different marks for the same essay answer. This shows low

    1. objectivity
    2. discrimination index
    3. usability
    4. difficulty index
    Answer

    A. objectivity

    Scorer-dependent marking means low objectivity.

  45. A test gives almost the same scores when a student takes it twice after a short gap. This indicates good

    1. reliability (test-retest)
    2. blueprint
    3. rubric
    4. validity of content only
    Answer

    A. reliability (test-retest)

    Stable results across repeated administration show reliability.

Page 1 of 1
‹
›