Nearly every public school teacher in the country is evaluated annually under a state system, a practice federal policy pushed toward standardization during the 2010s before the Every Student Succeeds Act returned design authority to states in 2015. The typical evaluation combines classroom observations scored on a rubric with evidence of student growth, and the exact weights vary by state and sometimes by district. For teachers and parents alike, understanding the machinery explains a lot about what schools prioritize.
What are the main components of a teacher evaluation?
Three pieces appear almost everywhere. Observations: an administrator or trained observer watches a lesson and scores it against a rubric. Student growth: measures of how much students improved, from state test-based value-added models in tested grades and subjects to district-set growth goals elsewhere. Additional evidence: student learning objectives, artifacts, professional responsibilities, sometimes student surveys. States set the recipe — for example, a common pattern is roughly half observation and half student growth, with local flexibility on the details.
What is a teaching rubric?
A rubric is a scoring framework that describes teaching behaviors at multiple performance levels, usually across three or four domains. Observers mark where the observed lesson falls on each dimension, producing a profile rather than a single grade. The two most widely used frameworks in U.S. schools are Danielson's Framework for Teaching and the Marzano evaluation models, each licensed and trained separately.
| Rubric domain (typical) | What observers look for |
|---|---|
| Planning and preparation | Coherent objectives, aligned activities, knowledge of students |
| Classroom environment | Routines, expectations, productive climate |
| Instruction | Questioning, engagement, feedback, pacing |
| Professional responsibilities | Reflection, communication with families, record accuracy |
How do observations actually work?
In most systems, a teacher receives one or two formal announced observations plus shorter unannounced walkthroughs each year. The observer scripts what students and the teacher do, then maps evidence to rubric language. A post-conference follows, ideally with specific evidence rather than adjectives. Problems with observation scoring are well documented: the Measures of Effective Teaching study, funded by the Gates Foundation and reported in 2013, found observers drift toward leniency and that scores rise when observers know the teacher, which is why many states require training and calibration for observers. Teachers rated proficient almost everywhere should not surprise anyone; the bell curve of observation scores is heavily compressed at the top.
What is a value-added model?
A value-added model, or VAM, is a statistical estimate of a teacher's contribution to student test-score growth, controlling for prior achievement and student characteristics. The idea is to isolate the teacher's effect from the students' starting point. The limitations matter as much as the idea: VAM estimates are noisy, fluctuate year to year, and are only computable in tested grades and subjects — roughly a quarter to a third of teachers. Multiple studies, including work published by the American Statistical Association in 2014, caution against high-stakes use of a single year of VAM data. Most states now cap how much test-based growth can count or allow districts to substitute locally set growth measures.
What are student learning objectives?
An SLO is a teacher-set growth goal: the teacher picks a learning target for a specific group of students, defines how it will be measured, and is evaluated on the results. SLOs spread widely because they extend growth measurement to untested subjects — art, physical education, early grades. Their weakness is comparability: an ambitious goal measured by a teacher-designed assessment is not equivalent across classrooms, and reviews of SLO implementation have found quality varies with the scrutiny applied to them.
What are the consequences attached to ratings?
Depends heavily on state law, and the trend since 2015 has been softening. Rating categories typically run from ineffective to exemplary across four levels. Consequences can include improvement plans, mentoring, timelines for earning tenure, and in the strictest states dismissal proceedings tied to consecutive low ratings. Due process for tenured teachers remains governed by state tenure law, which is why removing a teacher based on evaluations is slow in most states regardless of rating design. In many systems, the practical function of evaluation has shifted toward professional growth conversations rather than personnel removal.
How should a teacher read their own evaluation?
- Compare the scripted evidence to each rubric score; unscored evidence is grounds for a response.
- Check that the correct framework and weights were applied as listed in the state model.
- Note the timeline for the improvement or refinement plan if one is assigned.
- File a written response or appeal within the window the state or contract specifies.
What should parents take from all this?
Evaluation results are rarely public at the individual level; access varies by state, and many publish only school-level summaries. What parents can infer is system priorities: the rubric language a district uses describes what it considers good teaching, and the conferences attached to evaluations are where instructional coaching happens. A school with a functioning feedback loop — observations followed by substantive conversations — is usually doing the quiet work that improves instruction years before test scores move.
What do teachers say about the systems?
The practitioner critique focuses on time and alignment. Rubric language prizes certain visible behaviors — student-led discussion, posted objectives — that fit some subjects and grade levels more naturally than others, and kindergarten teachers, special education teachers, and career-technical instructors often report awkward fits. Announced observations invite a polished lesson that differs from the ordinary one, while unannounced visits can catch a review day and score it harshly. None of this invalidates the tool, but it explains why teachers push for multiple observations, trained observers, and evidence-based conferences over single-visit scores. States that shortened observation forms in recent years largely responded to exactly this feedback.
A structural point deserves attention: evaluation is a measurement system attached to a support system, and the two can drift apart. A district may rate teachers without funding coaching for those rated developing, which converts an improvement signal into a permanent label. When parents or board members ask about evaluation policy, the sharpest question is not how ratings are computed but what happens after a rating is assigned — what support, on what schedule, funded from which line.
Where is evaluation policy heading?
The direction since ESSA has been toward local flexibility, fewer high-stakes stakes attached to test-based measures, and renewed attention to retention. Several states have trimmed the student-growth share or paused evaluation consequences following pandemic-era disruptions, and legislative activity continues annually. Anyone reasoning from a specific state's rules should verify the current model with the state education agency, since weights and consequences have shifted repeatedly over the past decade.
For more context, read How to Read Your State School Report Card Without Getting Lost.
For more context, read homeschooling requirements by state.
For more context, read What MTSS and RTI Actually Mean for Struggling Students.
