Blog

Pass or fail: is testing a valid way to measure student progress?

Published January 14, 2023

By Matthew Lynch, July 10, 2017. Republished by Dreamlight High School for our students and families.

What if the measures we use to determine passing or failing grades are completely skewed? Is standardized testing, or any testing for that matter, the right way to determine student progress?

For obvious reasons, one of the first and most significant concerns for the application of standardized tests is that they are not consistent with the standards for fair and appropriate testing. Of course, educators must first define the standards themselves, and demonstrate them to be relevant. In this instance, we are referring to the standards for fair and appropriate testing as defined by the NRC Report, which says that “measurement validity refers to the extent to which evidence supports a proposed interpretation and use of test scores” for specific purposes.

For instance, a measurement validity of the reading section of the SAT I standard test would be assessed to have a reasonable validity for assessment of an individual’s reading comprehension skills, knowledge of grammar rules, and ability to make inferences from texts. The use of scores from this test to determine an individual’s preparedness for entry into a particular college program would also be reasonably good.

To go back to the more formal parameters, the general rule is that the internal structure of the test, the content of the test, the relationship of the test to other criteria, and the psychological processes and cognitive operations used by the examinee in responding to the test items must all support the purpose of the test.

Attribution of cause

A test assessing knowledge and skill should target the knowledge and skills specifically; looking, as well, to ensure that the knowledge and skills being assessed are those that have been obtained from appropriate instruction. In some instances, knowledge might depend on poor instruction or on factors that are unrelated to the skills under review. For instance, a student might score poorly on the SAT reading test because their teachers didn’t transfer the necessary knowledge and skill.

Another example would be that an individual might score badly on the SAT reading test not because they lack reading comprehension skills that the test intends to assess but because they have significant language barriers or because there are cultural differences that have some bearing on the test. A passage in American history that relies upon presupposed knowledge of American history or customs might be problematic and undermine the validity and fairness of test scores.

Disabilities can also factor as an issue for the attribution of cause. Several types of cognitive or even physical disabilities can undermine an individual’s performance in a testing scenario without appropriate interventions provided to support the student’s exceptionalities.

Opportunity to learn

In the context of K-12 assessments, the cause component also influences the extent to which students receive adequate opportunity to learn the material for the test. Adequate quality and quantity of instruction become important, as does the alignment of test content and curriculum.

Students need adequate opportunity within the testing scenarios to demonstrate their knowledge. If tests contain irrelevant language or content, students may not have adequate opportunity to perform and test developers will have compromised the fairness and relevance of the test.

Furthermore, many of the criteria for fairness in testing standards overlap with attribution of cause: the investigation of bias and differential item functioning, determining whether construct-irrelevant variance differentially affects different groups of examinees, and equal treatment during the testing process.

Curricular validity relates to the alignment between test content and the curriculum taught in class. Chapter 13 of the Standards determines that “there should be evidence that the test adequately covers only the specific or generalized content and skills that students have had an opportunity to learn.”

Who is accountable?

Certain policies within the K-12 setting make high-stakes student decisions dependent upon evidence that the student has had the educational experience and opportunity to acquire relevant knowledge and skill. Where students have lacked sufficient opportunity to acquire desired skills, they may not meet the criteria for grade promotion or graduation.

At the same time, it is hardly fair that the student be held accountable for the deficit in their learning. At what point do we say: this portion of education is the responsibility of the schools, of the system and the stakeholders, not just the individual student?

The effectiveness of treatment is the final component of the fair and appropriate test criteria, relating to whether test scores lead to consequences that are educationally beneficial in a given context. Accountability plays a part here, too: it is inappropriate to use tests to make placements that are not educationally beneficial.

When tests are used in placement decisions, they must be fair and appropriate. Students must be “better off in the setting in which they are placed than they would be in a different available setting.” With all of these factors in mind, can testing ever truly be trusted as a placement option for students?