EdReports, the nonprofit that independently reviews K-12 instructional materials against standards-alignment criteria, publishes ratings precisely because districts otherwise have almost no neutral way to judge whether a curriculum's first rough year reflects a flawed program or a rushed rollout. Curriculum directors who conflate the two tend to either scrap materials that needed one more year of training, or keep materials that were never going to work regardless of implementation quality.
This is a framework for making that distinction, not a verdict on any specific publisher's materials. Every district's evidence is different, and this piece does not substitute for a district's own data.
What Separates an Implementation Problem From a Design Problem?
Implementation problems show up as inconsistency: some classrooms with strong outcomes and others with weak outcomes, teacher survey data that clusters around insufficient training or pacing confusion rather than the materials themselves, and improvement over the course of the year as staff got more familiar with the sequence. Design problems show up as consistency in the wrong direction: outcomes flat or declining across most classrooms regardless of teacher experience, and specific standards or skill areas the materials simply do not address adequately, verifiable against an independent alignment review such as an EdReports rating.
A single year of data rarely settles the question on its own. The distinction usually requires triangulating implementation fidelity — was the curriculum taught as designed, with adequate training and pacing — against outcome data disaggregated by classroom, not just by school.
What Evidence Should a Renewal Decision Actually Rest On?
A defensible renewal review draws on multiple sources, not a single test score trend: interim and end-of-year assessment data by classroom and student subgroup; teacher survey and focus-group data on pacing, training adequacy, and material quality specifically; classroom observation data on whether the curriculum was implemented with fidelity or adapted heavily; and an independent alignment rating, where one exists, as a check against the district's own read of quality.
Curriculum offices that skip the fidelity check most often end up making a design judgment on implementation data — concluding materials failed when the real problem was that teachers received two days of training for a curriculum that needed ten.
What Does a Fair Second-Year Trial Look Like?
If year one shows implementation gaps rather than design gaps, a second year with a defined improvement plan is the more defensible call: targeted professional development on the specific pacing or content areas teachers flagged, coaching support concentrated on classrooms furthest behind, and pre-committed metrics for what a successful second year looks like, decided before the year starts rather than after.
What makes this defensible to a school board or a skeptical staff is the pre-commitment. A district that sets its renewal criteria after seeing the data invites the reasonable objection that the bar moved to fit the outcome.
When Is Replacement the Right Call?
Replacement is justified when the alignment review itself is weak — an independent rating flags gaps in standards coverage that no amount of training fixes — or when a second year with a real improvement plan still shows flat results across most classrooms, not just the ones with the least experienced teachers. It is also justified when teacher attrition or morale data tied specifically to the curriculum threatens broader retention, a cost that does not always show up in test-score data but is real.
Replacement decided after one implementation-limited year, without an improvement plan, is the pattern curriculum researchers most often flag as wasteful: it burns the sunk cost of training and materials adoption while never testing whether the curriculum itself, properly implemented, would have worked.
How Should a District Communicate the Decision?
Whichever way the decision goes, publishing the evidence base — the disaggregated outcome data, the fidelity findings, the alignment rating — protects the district from the perception that the call was political rather than evidentiary, and gives the next curriculum office a documented precedent to work from rather than starting the argument over from scratch.




