HSCredLog in

Research

The argument, with receipts.

Explore the evidence ↓

The papers and findings behind a credential that can be compared without flattening a student into a test score.

What the evidence says.

Written authenticity has collapsed.

Generative AI produces fluent, on-prompt application essays at no cost. Half of admissions officers view AI use in essays unfavorably; only fourteen percent view it favorably.

Kaplan, 2025

Institutions are deciding at record volume with no trusted signal.

First-year applications passed ten million for the first time in 2024-25, up 161% in a decade, while the offices reading them have not grown to match.

Common App, 2025

Detection cannot fix it.

Leading detectors are biased against non-native English writers, their accuracy collapses under simple paraphrasing, and OpenAI withdrew its own detector for low accuracy.

Liang et al., 2023; OpenAI, 2023; Perkins et al., 2024

Supervised, defended work is what survives.

A credential whose provenance is established through supervision and live revision, not inferred after submission, is the one generative AI cannot counterfeit.

Dawson, 2021; Dawson et al., 2024; Eaton, 2023

Every claim, out in the open.

Transparency is a core commitment at HSCred, so the full evidence base is here in public. You do not have to be a researcher to use it: every key claim is laid out with the studies behind it, each tagged by how strong the evidence is and paired with its honest limitation, so you can judge for yourself whether the argument holds, and see exactly where the questions are still open.

98 sources13 themesTagged by study type & limitation
Filter by method
What the tiers mean

Direct test. Experiments and randomized trials that test a claim head-on.

Field study. Validity, program, and quasi-experimental studies from real schools and admissions offices.

Synthesis & data. Meta-analyses, research reviews, and datasets that pool many studies or measure the whole field.

Theory & source. Frameworks that define the terms, and primary documents, policy, rulings, first-party records.

Tiers describe the kind of evidence, not a ranking, a meta-analysis of 69 studies and a single experiment carry weight in different ways. Use them to filter, not to score.

Showing 98 of 98
01

The authenticity collapse: AI and the failure of detection

Take-home written work can no longer prove its own authorship, and detection cannot repair that.

Generative AI writes fluent, on-prompt essays for free, and the tools sold to catch it are unreliable and biased, OpenAI withdrew its own detector for low accuracy. The answer isn't better detection but provenance: work supervised and defended as it is made, so authorship never has to be inferred after the fact.

Report / data

Admissions officers distrust AI-assisted essays: half view the practice unfavorably, only 14% favorably.

Kaplan. (2025). College admissions officers survey: Applicants' use of AI in essays.

Limitation An industry survey of perceptions, not an audit of essay authorship.

Experiment

Seven detectors flagged most TOEFL essays by non-native writers as AI-generated: detection is unreliable and unfair.

Liang, W., et al. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7).

Limitation Evaluated 2023-era detectors; accuracy shifts as models change.

Experiment

Detector accuracy collapses under simple paraphrasing; the authors advise against using detectors to determine violations.

Perkins, M., et al. (2024). Simple techniques to bypass GenAI text detectors. Int. J. Educational Technology in Higher Education, 21.

Limitation Tests specific detectors and evasion tactics; a moving target.

Primary document

The model maker's own detector identified only a minority of AI text and was withdrawn for low accuracy.

OpenAI. (2023). New AI classifier for indicating AI-written text.

Limitation A vendor product note, not an independent evaluation.

Theory

Establishes provenance through supervision and authentication rather than after-the-fact inference.

Dawson, P. (2021). Defending assessment security in a digital world. Routledge.

Limitation Foundational framing, not an empirical result.

Theory

Reframes the AI problem as a threat to assessment validity rather than to honesty.

Dawson, P., Bearman, M., Dollinger, M., & Boud, D. (2024). Validity matters more than cheating. Assessment & Evaluation in Higher Education, 49(7).

Limitation Argument rather than data.

Theory

Introduces the 4Ps (product, process, performance, practice); a written product is the weakest proxy on its own, a limitation AI compounds.

Fawns, T., Boud, D., & Dawson, P. (2026). Identifying what our students have learned: A framework for practical assessment validation. Assessment & Evaluation in Higher Education.

Limitation A conceptual framework, not an empirical study.

Theory

Hybrid human-and-machine authorship makes separating the two in take-home text impractical in principle.

Eaton, S. E. (2023). Postplagiarism. Int. J. for Educational Integrity, 19, Article 23.

Limitation Conceptual, and used as such.

02

Evidence over labels: measurement and predictive validity

Selection rewards predictive validity; high-inference labels do not travel between institutions, and richer validated evidence improves prediction.

Selection rewards signals that predict success, and high-school GPA does, while a letter grade carries no stable meaning across 27,000 schools. Adding validated, richer evidence of what a student can actually do improves prediction and narrows group differences.

Theory

The argument-based framework for validity: more ambitious claims require more evidence.

Kane, M. T. (2013). Validating the interpretations and uses of test scores. J. Educational Measurement, 50(1).

Limitation A validity framework, not empirical evidence.

Research review

A century of evidence that grades blend achievement with unrelated factors and carry no stable cross-school meaning.

Brookhart, S. M., et al. (2016). A century of grading research. Review of Educational Research, 86(4).

Limitation A synthesis of older studies; it documents the problem, not a fix.

Validity study

High-school GPA strongly and consistently predicts college completion; the ACT's relationship is weak and varies across schools.

Allensworth, E. M., & Clark, K. (2020). High school GPAs and ACT scores as predictors of college completion. Educational Researcher, 49(3).

Limitation Observational cohort study, not a randomized comparison.

Validity study

HS GPA predicts four-year outcomes better than the SAT and correlates far less with socioeconomic status.

Geiser, S., & Santelices, M. V. (2007). Validity of high-school grades in predicting student success beyond the freshman year. CSHE, UC Berkeley.

Limitation A UC sample and a working paper, so paired with Allensworth & Clark.

Validity study

Analytical, creative, and practical assessments added predictive validity beyond the SAT and HS GPA, while reducing group differences.

Sternberg, R. J., & The Rainbow Project Collaborators. (2006). The Rainbow Project. Intelligence, 34(4).

Limitation College Board-funded, with limited independent replication.

Theory

The Nobel-recognized result that a credible signal must be costly to fake.

Spence, M. (1973). Job market signaling. Quarterly J. of Economics, 87(3).

Limitation Theory, applied by analogy.

Program evaluation

Competency-based pilots were heterogeneous and explicitly non-causal.

Steele, J. L., et al. (2014). Competency-based education in three pilot programs. RAND.

Limitation Cited for its limits, not for a positive result.

03

The limits of standardized scores and grades

Before proposing a new standard, the brief documents why the existing ones underdetermine capability.

Before proposing a new standard, the brief documents why the current ones underdetermine capability: grades inflated for a decade while test performance stalled, and scores track family income closely enough to read as an opportunity map as much as an ability one.

Theory

A measurement scholar's account of how scores under high-stakes accountability inflate and narrow instruction.

Koretz, D. (2017). The testing charade. University of Chicago Press.

Limitation A synthesis, not a single study.

Report / data

The test maker's own data: average grades rose for a decade while test performance stagnated.

ACT. (2022). Grade inflation continues to grow in the past decade. Report R2134.

Limitation Notable because the source has no incentive to undercut grades.

Correlational

Family income substantially influences SAT performance, largest for low-income and Black students.

Dixon-Roman, E. J., Everson, H. T., & McArdle, J. J. (2013). Race, poverty and SAT scores. Teachers College Record, 115(4).

Limitation Cited honestly as substantial but secondary to prior achievement.

Quasi-experiment

Documents how admissions advantages track family income at the top of the distribution.

Chetty, R., Deming, D. J., & Friedman, J. N. (2023). Diversifying society's leaders? NBER Working Paper 31492.

Limitation A working paper, not yet peer reviewed.

Report / data

An institutional explainer summarizing score-gap data.

Georgetown University. (2023). SAT score gaps reveal deeper inequality.

Limitation Secondary; the brief attributes framing to Georgetown and data to its primary source.

Report / data

A large multi-institution study: dropping a common signal left institutions inferring from noisier proxies.

Syverson, S. T., Franks, V. W., & Hiss, W. C. (2018). Defining access: How test-optional works. NACAC.

Limitation Professional-association research.

Theory

Establishes the opportunity-gap framing: a measurement problem before it is an equity problem.

Carter, P. L., & Welner, K. G. (Eds.). (2013). Closing the opportunity gap. Oxford University Press.

Limitation Synthesis, cited for framing.

04

The empirical record of performance assessment

The outcome evidence that performance-based assessment predicts success: consistent and multi-site, but largely non-randomized.

Where students are admitted or graduated on defended, performance-based work, they persist and succeed at least as well as their peers, often with the largest gains for lower-achieving students. The evidence is consistent across many sites but largely quasi-experimental, not randomized.

Program evaluation

CUNY / NY Performance Standards Consortium: students with lower SATs earned higher first-semester grades and persisted at higher rates.

Fine, M., & Pryiomka, K. (2020). Assessing college readiness through authentic student work. Learning Policy Institute.

Limitation Observational, not randomized.

Quasi-experiment

NH PACE schools showed comparable to modestly improved outcomes, with larger gains for lower-achieving students.

Evans, C. M. (2019). Effects of New Hampshire's innovative assessment system. Education Policy Analysis Archives, 27.

Limitation Non-randomized but peer reviewed.

Quasi-experiment

A multi-year follow-up showing durable no-harm from substituting performance assessment for standardized testing.

Perez, A. L., & Evans, C. (2023). Keeping up the PACE. Applied Measurement in Education, 36(2).

Limitation Non-randomized; extends Evans (2019).

Quasi-experiment

Across ~20,000 students in 19 deeper-learning schools: higher graduation, test scores, and college enrollment than matched peers.

Zeiser, K. L., et al. (2014). Evidence of deeper learning outcomes. American Institutes for Research.

Limitation Matched design, not randomized.

Quasi-experiment

The longer-run follow-up in which several college advantages were no longer statistically significant.

American Institutes for Research. (2022). The study of deeper learning: College outcomes (Report 7).

Limitation Included deliberately for honesty about mixed long-run results.

Correlational

A supervised, defended long-form research task predicted later college GPA.

Inkelas, K. K., et al. (2013). Benefits of the IB extended essay. International Baccalaureate Organization.

Limitation Sponsored and observational.

Validity study

Combining multiple forms of evidence improves prediction of college readiness beyond a single test.

Gaertner, M. N., & McClarty, K. L. (2015). Performance, perseverance, and the full picture of college readiness. Educational Measurement, 34(2).

Limitation A review, not a single controlled trial.

Validity study

Summarizes the Rainbow and Kaleidoscope projects: performance and creative-practical assessments add predictive validity beyond the SAT.

Sternberg, R. J. (2010). College admissions for the 21st century. Harvard University Press.

Limitation One author's synthesis of his own projects; corroboration only.

Report / data

A policy synthesis of performance-assessment models and their adoption.

Guha, R., et al. (2018). The promise of performance assessments. Learning Policy Institute.

Limitation Advocacy-adjacent institute, weighted as synthesis.

05

Inter-rater reliability and rubric quality

The reliability objection, met head-on: the historical failure point, and the supports that raise agreement.

The oldest objection to portfolios is that two readers will not agree. The record shows that is a solved engineering problem: shared rubrics, rater training, calibration, and moderation raise agreement to levels that support comparison.

Program evaluation

The classic demonstration that portfolio scoring across many schools had inter-rater agreement too low to support comparison.

Koretz, D., Stecher, B., Klein, S., & McCaffrey, D. (1994). The Vermont portfolio assessment program. Educational Measurement, 13(3).

Limitation The cautionary baseline the brief concedes rather than hides.

Research review

Well-designed rubrics with training raise the reliability of judgment and direct it toward substance.

Jonsson, A., & Svingby, G. (2007). The use of scoring rubrics. Educational Research Review, 2(2).

Limitation A 2007 review; the pooled studies vary in rigor.

Research review

Rubric use improves the consistency and transparency of scoring in higher education.

Reddy, Y. M., & Andrade, H. (2010). A review of rubric use in higher education. Assessment & Evaluation in Higher Education, 35(4).

Limitation A narrative review focused on higher education.

Methodology

The formal framework for partitioning sources of error in performance scores: reliability engineered, not asserted.

Brennan, R. L. (2000). Performance assessments from the perspective of generalizability theory. Applied Psychological Measurement, 24(4).

Limitation Theory and method.

Methodology

The standard method for detecting a persistently lenient or severe rater, behind the flag-and-retrain mechanism.

Myford, C. M., & Wolfe, E. W. (2003). Detecting and measuring rater effects using many-facet Rasch measurement. J. Applied Measurement, 4(4).

Limitation Method, not outcome.

Report / data

Documents the Consortium's shared rubrics, calibration, and moderation across schools.

Cook, A., & Tashlik, P. (2017). Building a system of assessment: The NY Performance Standards Consortium. In Beyond Testing. Teachers College Press.

Limitation Descriptive account of practice, not an independent audit.

Primary document

The Consortium's own record of its scale and history.

New York Performance Standards Consortium. (n.d.). History of the Consortium.

Limitation First-party, used only for descriptive facts.

Report / data

Resistance to gaming by production values is a design property of high-quality performance assessment.

Badrinarayan, A. (2022). Performance assessments in college admission. Learning Policy Institute.

Limitation Policy report.

Report / data

A rubric-scored portfolio where a reader rewarded a practice-session video over a polished final; jargon can distract from thinking.

Willis, L., & Martinez, M. R. (2022). Authentic student work in college admissions: Ross School of Business. Learning Policy Institute.

Limitation A single documented case, illustrative.

06

Content mastery as a precondition: the science of learning

The answer to 'portfolios bypass rigor': demonstrated content mastery is required before project work. The most causally robust theme.

The most causally robust theme. Critical thinking is bound to domain knowledge, novices need schemas before open inquiry, and mastery learning reliably raises achievement, most for the weakest students. Project work rests on content mastery rather than bypassing it.

Research review

Critical thinking is bound to domain knowledge, not a content-free skill.

Willingham, D. T. (2007). Critical thinking: Why is it so hard to teach? American Educator, 31(2).

Limitation A magazine-format review of research.

Research review

The 'generic skills' schools prize are in practice applications of domain-specific knowledge.

Tricot, A., & Sweller, J. (2014). Domain-specific knowledge and why teaching generic skills does not work. Educational Psychology Review, 26(2).

Limitation Peer-reviewed synthesis.

Research review

Novices set loose on inquiry without prerequisite knowledge are overloaded rather than enabled.

Kirschner, P. A., Sweller, J., & Clark, R. E. (2006). Why minimal guidance during instruction does not work. Educational Psychologist, 41(2).

Limitation Contested in the field; cited for the prerequisite-knowledge point.

Experiment

The founding cognitive-load experiments: problem solving without schemas burdens working memory.

Sweller, J. (1988). Cognitive load during problem solving. Cognitive Science, 12(2).

Limitation Early laboratory studies; a mechanism, not an outcome.

Experiment

Expert performance rests on organized domain knowledge, not general ability.

Chi, M. T. H., Feltovich, P. J., & Glaser, R. (1981). Categorization of physics problems by experts and novices. Cognitive Science, 5(2).

Limitation Small expert-novice samples; a mechanism, not an outcome study.

Theory

Introduced mastery learning: command of prerequisites before advancing.

Bloom, B. S. (1968). Learning for mastery. Evaluation Comment, 1(2).

Limitation The origin statement; evidence comes from the meta-analyses below.

Meta-analysis

Mastery learning raised achievement by ~0.5 SD on average, with larger gains for weaker students.

Kulik, C.-L. C., Kulik, J. A., & Bangert-Drowns, R. L. (1990). Effectiveness of mastery learning programs: A meta-analysis. Review of Educational Research, 60(2).

Limitation Effects were larger on locally made than standardized tests.

Experiment

Actively recalling material beats restudy for long-term retention.

Roediger, H. L., & Karpicke, J. D. (2006). Test-enhanced learning. Psychological Science, 17(3).

Limitation A laboratory paradigm; classroom transfer varies.

Experiment

Structured student autonomy raises motivation and engagement.

Cullen, S., & Oppenheimer, D. M. (2024). Choosing to learn: student autonomy in higher education. Science Advances, 10(29).

Limitation A single recent experiment; short-term outcomes.

Theory

Intrinsic motivation grows when learners experience competence and autonomy.

Ryan, R. M., & Deci, E. L. (2000). Self-determination theory. American Psychologist, 55(1).

Limitation A theory, not a single empirical result.

07

Production as learning and the video medium

Why host produced video rather than written work, on learning grounds, including a direct comparison of video against writing.

Explaining your work is itself a way of learning it, and the medium matters: in a head-to-head experiment, explaining on video improved learning while explaining in writing did not. Producing a defense, especially for an audience, is where the understanding gets built.

Research review

Generative strategies, among them self-explaining and teaching, deepen understanding relative to passive study.

Fiorella, L., & Mayer, R. E. (2016). Eight ways to promote generative learning. Educational Psychology Review, 28(4).

Limitation A review of strategies; effects vary by setting.

Meta-analysis

Across 69 effect sizes, prompting self-explanation produced a moderate benefit (g ~ 0.55).

Bisra, K., et al. (2018). Inducing self-explanation: A meta-analysis. Educational Psychology Review, 30(3).

Limitation Benefits vary with prompt quality and domain.

Experiment

The decisive comparison: explaining on video improved learning over restudy while explaining in writing did not.

Hoogerheide, V., et al. (2016). Gaining from explaining: learning improves from explaining on video, not from writing. Contemporary Educational Psychology, 44-45.

Limitation Short-term learning with student samples; not an admissions study.

Experiment

Students who created an instructional video outperformed those who restudied, and enjoyed it more.

Hoogerheide, V., et al. (2019). Generating an instructional video as homework. Learning and Instruction, 64.

Limitation A controlled study with student samples, not classroom-wide.

Experiment

Studying while expecting to teach, rather than to be tested, improved recall and organization.

Nestojko, J. F., et al. (2014). Expecting to teach enhances learning. Memory & Cognition, 42(7).

Limitation A laboratory free-recall study.

Experiment

Actually teaching raised comprehension; preparing to teach by itself did not.

Fiorella, L., & Mayer, R. E. (2013). The relative benefits of learning by teaching and teaching expectancy. Contemporary Educational Psychology, 38(4).

Limitation A single experiment; may not generalize.

08

Project-based and deeper learning

Meta-analyses and randomized trials on achievement under project-based methods, kept distinct from the admissions-prediction evidence.

Across meta-analyses and multi-district randomized trials, project-based instruction beats traditional teaching on achievement, including on standardized tests and for low-income students. This is the theme with the most randomized evidence behind it.

Meta-analysis

Project-based learning produced a moderate-to-large achievement effect (~d = 0.65) against traditional instruction.

Zhang, L., & Ma, Y. (2023). Impact of project-based learning: A meta-analysis. Frontiers in Psychology, 14.

Limitation Pooled studies of varying quality.

Meta-analysis

Project-based learning outperformed traditional instruction on achievement (~d = 0.71), moderated by subject and design.

Chen, C.-H., & Yang, Y.-C. (2019). Revisiting the effects of project-based learning: A meta-analysis. Educational Research Review, 26.

Limitation Effects moderated by subject and design; study quality varies.

Experiment

In a multi-district randomized trial, students in project-based AP courses earned credit-qualifying scores at higher rates, including low-income students.

Saavedra, A. R., et al. (2022). The impact of project-based learning on AP exam performance. Educational Evaluation and Policy Analysis, 44(4).

Limitation A multi-district randomized trial, but only two AP subjects.

Experiment

Randomized project-based units raised 2nd graders' social-studies and informational-reading achievement in low-income schools.

Duke, N. K., et al. (2021). Putting PjBL to the test. American Educational Research Journal, 58(1).

Limitation Effects held for achievement, not writing or motivation.

Experiment

A randomized project-based science program raised 3rd graders' standardized science scores by ~0.28 SD.

Krajcik, J., et al. (2023). Assessing the effect of project-based learning on science learning. American Educational Research Journal, 60(1).

Limitation Independent replication; an elementary sample.

Theory

The learning-sciences account of why project-based methods support deep learning.

Krajcik, J. S., & Blumenfeld, P. C. (2006). Project-based learning. In The Cambridge Handbook of the Learning Sciences.

Limitation Cited for mechanism; effect sizes carried by the meta-analyses.

09

Journalism across the curriculum

The on-ramp argument: the writing-to-learn tradition, making thinking visible, and scholastic-journalism outcomes.

Writing is a mode of thinking, not just a record of it, and a real audience sharpens both the cognition and the motivation. Scholastic journalism, controlling for prior differences, still predicts stronger English scores and later civic participation, most for lower-income students.

Theory

The founding claim that writing is a mode of learning, not merely a record of it.

Emig, J. (1977). Writing as a mode of learning. College Composition and Communication, 28(2).

Limitation A 1977 theoretical essay, not empirical evidence.

Meta-analysis

Writing-to-learn produced small but consistent achievement gains, larger when students reflected on their own understanding.

Bangert-Drowns, R. L., Hurley, M. M., & Wilkinson, B. (2004). Effects of writing-to-learn interventions: A meta-analysis. Review of Educational Research, 74(1).

Limitation Average gains are small and task-dependent.

Empirical study

Different writing tasks engage different kinds of thinking and shape content understanding.

Langer, J. A., & Applebee, A. N. (1987). How writing shapes thinking. NCTE.

Limitation Mixed methods, not randomized.

Theory

Learning by making expert thinking observable.

Collins, A., Brown, J. S., & Holum, A. (1991). Cognitive apprenticeship: Making thinking visible. American Educator, 15(3).

Limitation A theoretical account, not an outcome study.

Research review

A real audience shapes both the cognition and the motivation of writing.

Magnifico, A. M. (2010). Writing for whom? Cognition, motivation, and a writer's audience. Educational Psychologist, 45(3).

Limitation A review, not a controlled study.

Theory

Defines authentic intellectual work as inquiry yielding products with value beyond school.

Newmann, F. M., & Wehlage, G. G. (1993). Five standards of authentic instruction. Educational Leadership, 50(7).

Limitation A framework, not empirical evidence.

Theory

Assessment is authentic when it examines real performance on worthy tasks rather than proxies.

Wiggins, G. (1990). The case for authentic assessment. Practical Assessment, Research, and Evaluation, 2.

Limitation A definitional argument, not evidence.

Correlational

After controlling for prior characteristics, more high-school journalism still predicted higher standardized English scores.

Bobkowski, P. S., & Cavanah, S. B. (2019). When 'journalism kids' do better. Journalism & Mass Communication Educator, 74(4).

Limitation Controlled for prior traits, but still correlational.

Correlational

Associates high-school journalism with later voting and volunteering, more so for lower-income students.

Bobkowski, P. S., & Miller, P. R. (2016). Civic implications of secondary school journalism. Journalism & Mass Communication Quarterly, 93(3).

Limitation National data, association only.

10

The economy that rewards this work

Why the labor market rewards investigation, synthesis, and communication over routine recall.

The labor market increasingly pays for exactly what routine schooling under-measures: non-routine problem-solving, synthesis, and communication, especially where analytical and social skill combine. Recall is the part machines already do cheaply.

Empirical study

Computers substitute for routine tasks and complement non-routine problem solving and communication.

Autor, D. H., Levy, F., & Murnane, R. J. (2003). The skill content of recent technological change. Quarterly J. of Economics, 118(4).

Limitation A decomposition of task demand over time, not an experiment.

Empirical study

Rising labor-market returns to roles combining analytical and social skill.

Deming, D. J. (2017). The growing importance of social skills in the labor market. Quarterly J. of Economics, 132(4).

Limitation Observational, peer reviewed.

Theory

A wealth of information creates a poverty of attention.

Simon, H. A. (1971). Designing organizations for an information-rich world.

Limitation Conceptual.

11

The feasibility constraint: teacher time

Any new assessment must not add to teacher workload. The brief's answer: the artifact is the record.

Any new assessment has to fit teachers who already work far beyond contract, much of it on documentation. The brief's answer is that the artifact is the record, the defended work is the evidence, so assessment stops being extra paperwork.

Report / data

Teachers work far beyond contracted hours (~53 vs 44), much of the excess on documentation rather than instruction.

Doan, S., Steiner, E. D., & Pandey, R. (2024). Teacher well-being and intentions to leave in 2024. RAND.

Limitation National survey, self-reported.

Report / data

Educators named the need for more time as the most prevalent implementation theme; the state later made the mandate optional.

Johnson, A., & Stump, E. (2018). Proficiency-based diploma systems in Maine. Maine Education Policy Research Institute.

Limitation A single-state perception study.

12

Institutions, standards, policy, and access

The institutional case: standards bodies, federal and state policy, access data, and demand signals.

The institutional ground is already shifting: federal law now bars race-conscious admissions and explicitly permits performance-based assessment, states are adopting Portrait-of-a-Graduate standards, and access data shows advanced coursework is unevenly available. Even selective schools reinstating tests frame them as the least-bad common signal they have, not the one they would choose.

Legal authority

The decision prohibiting race-conscious admissions, creating the need for race-neutral talent identification.

Students for Fair Admissions v. President and Fellows of Harvard College, 600 U.S. 181 (2023).

Limitation Primary legal source.

Legal authority

The Innovative Assessment Demonstration Authority explicitly enumerates performance-based assessment as permissible.

Every Student Succeeds Act of 2015, 20 U.S.C. Sec. 6364.

Limitation Primary statutory source.

Report / data

The official IES evaluation of the authority's early implementation.

Troppe, P., et al. (2023). Evaluating the federal Innovative Assessment Demonstration Authority (NCEE 2023004). IES.

Limitation Independent federal evaluation.

Primary document

The formal adoption of the Portrait of a Graduate framework.

New York State Education Department. (2025). Board of Regents adopt New York State Portrait of a Graduate.

Limitation First-party state policy.

Primary document

The Portrait's six attributes in the state's own words.

New York State Education Department. (2025). NYS Portrait of a Graduate [six attributes].

Limitation First-party state policy.

Report / data

Federal audit tying the advanced-course gap to school poverty and size.

U.S. Government Accountability Office. (2018). Public high schools with more students in poverty offer fewer academic offerings (GAO-19-8).

Limitation Independent government data.

Report / data

National civil-rights data: advanced coursework is less available in schools serving mostly Black and Hispanic students.

U.S. Dept. of Education, OCR. (2024). Student access to math, science, and computer science (2020-21 CRDC).

Limitation Census-level federal data.

Report / data

The test maker's own data on AP-course availability by group.

College Board. (2025). AP national and state data.

Limitation First-party data; interest noted.

Report / data

The current count of test-optional and test-free institutions.

FairTest. (2025). Overwhelming majority of U.S. colleges remain test-optional or test-blind for fall 2026.

Limitation Advocacy organization; figure attributed to FairTest.

Primary document

A leading institution treats the test as the most equitable common signal it currently has, while stating it would adopt a better one.

MIT Admissions. (2022). We are reinstating our SAT/ACT requirement.

Limitation An admissions-office position statement, not independent evidence.

Primary document

Reinstatement resting on faculty analysis that scores, read against local norms, surfaced disadvantaged high-achievers.

Dartmouth College. (2024). Reactivating the SAT/ACT requirement.

Limitation One institution's account of its own internal study.

Primary document

A test-flexible requirement (AP or IB may substitute), citing findings that test-optional disadvantaged lower-income applicants.

Yale University. (2024). Yale announces new test-flexible admissions policy.

Limitation The reinstatement trend is not uniformly an SAT/ACT mandate.

Primary document

A committee found academic outcomes correlate with scores across subgroups and that scores read in context can widen access.

Brown University. (2024). Brown to reinstate test requirement.

Limitation A committee summary specific to Brown's pool.

Primary document

Frames the test as the least-biased common signal currently available, citing Chetty et al.

Harvard University. (2024). Harvard announces return to required testing.

Limitation A news article summarizing research.

Primary document

A top engineering institution's optional Maker Portfolio, reviewed by faculty and alumni with maker expertise.

Massachusetts Institute of Technology. (2026). First-year applicants: Creative portfolios.

Limitation An optional supplement; cross-disciplinary demand for authentic work.

Primary document

Every advancing applicant attends a Candidates' Weekend whose design challenge is scored by the Olin community.

Franklin W. Olin College of Engineering. (2026). Admission process.

Limitation A required performance assessment, not an add-on.

Primary document

An alternative route: admission on long analytical essays graded by Bard faculty like coursework.

Bard College. (2025). The Bard Entrance Examination.

Limitation The clearest case of admission on academic merit alone.

Report / data

The registrar profession's initiative for richer, verifiable, portable records of achievement.

American Association of Collegiate Registrars and Admissions Officers, & NASPA. (2018). Comprehensive Learner Record.

Limitation Establishes professional demand for the container.

13

Classroom implementation: differentiation and mathematics

Flexible within-class grouping lets one classroom serve students at different stages, and mathematics is where the discipline's own standards ask for depth.

At the classroom level, flexible within-class grouping raises achievement at every level where fixed tracking does not, and mathematics' own professional standards call for the productive struggle and high-demand tasks that defended project work is built to require.

Meta-analysis

Within-class and cross-grade flexible grouping raised achievement at all levels, while fixed between-class tracking did not.

Steenbergen-Hu, S., Makel, M. C., & Olszewski-Kubilius, P. (2016). One hundred years of research on ability grouping and acceleration. Review of Educational Research, 86(4).

Limitation A synthesis of syntheses; only as strong as the studies it pools.

Primary document

The field's professional standard naming productive struggle and high-cognitive-demand tasks as core practices.

National Council of Teachers of Mathematics. (2014). Principles to actions.

Limitation Authoritative consensus, cited for what the discipline calls rigorous.

Research review

Students' productive struggle with important mathematics is a condition for understanding it.

Hiebert, J., & Grouws, D. A. (2007). The effects of classroom mathematics teaching on students' learning. In Second Handbook of Research on Mathematics Teaching and Learning.

Limitation A handbook synthesis, widely cited.

Experiment

Students who grappled with novel problems before instruction developed deeper conceptual understanding.

Kapur, M. (2014). Productive failure in learning math. Cognitive Science, 38(5).

Limitation Direct experimental support; modest sample sizes.

Openings for new research

Where the evidence is still thin

The most useful thing an evidence base can do is show its own edges. These are the places where our claims outrun the strongest available studies, each one a question worth a new paper.

Randomized evidence for defended credentials

The performance-assessment record is consistent but mostly quasi-experimental. A preregistered randomized trial comparing a supervised, defended video credential against a written essay, on later college GPA and persistence, would move this from correlation to cause.

Detector reliability is a moving target

The detection studies evaluate 2023–2024 models, and each new model generation resets the finding. A living benchmark that re-tests detectors as models ship would keep the no-detection claim current instead of dated.

Inter-rater reliability on HSCred's own instrument

The rubric-reliability evidence comes from other programs' portfolios. Publishing HSCred's own generalizability coefficients and rater-drift flags, on its own defended work, would test the reliability claim on the actual instrument rather than by analogy.

Video vs. writing, replicated at admissions stakes

The decisive video-beats-writing result is a laboratory learning study. Whether a produced video defense predicts college readiness better than an essay, in a real applicant pool, has not been tested.

Who a defended-work credential leaves out

Access data shows advanced coursework is unevenly available. Whether a defended-work credential widens access or re-narrows it depends on who has the supervision and equipment to produce one, an equity question the current evidence does not answer.

Teacher time, measured rather than argued

The feasibility case rests on “the artifact is the record” as an argument, not a time study. A workload measurement of teachers running defended-work assessment against traditional grading would confirm or refute it.