Back to Principal's Playbook

T-TESS 2 Calibration - How to Recalibrate Your Appraisers

By Observation Copilot Team

T-TESS 2 calibration is the work of getting your appraisal team to agree on what the new indicators mean before anyone rates a teacher on them. Recertification training teaches the rubric. It does not produce agreement. Two changes break whatever agreement your team already had: named target indicators inside each dimension, and a rule that a descriptor must generally be met at least 80% of the time. Both demand a norming session your team has not run yet.

This is the third post in a series on the refreshed Texas rubric. Start with what actually changed and then when it happens. This post is the execution piece: what a leadership team does in a room, on a specific afternoon, to keep ratings defensible across a version change.

Why does a new rubric version break appraiser calibration?

Because your team was normed on language that no longer exists. Whatever shared interpretation of Proficient your assistant principals built over three years was built against 16 dimensions and their descriptors. T-TESS 2 reorganizes those dimensions, renames two domains, and inserts named indicators underneath each one. The agreement does not transfer. It was agreement about specific sentences.

Rater judgment does not hold still even when the rubric does. The most thorough study of the effect, published in Educational and Psychological Measurement, tracked observers across roughly two years of scoring and found that raters became steadily more severe over time, dropping 0.7 to 1.3 points on a seven-point scale. More importantly, variability among raters increased. Evaluators left alone drift apart, not together.

So a version change is not a fresh start. It is a reset applied to a team that was already dispersing. Our general calibration protocol covers the routine version of this work. What follows is the part specific to changing rubrics underneath a team mid-career.

What does the T-TESS 2 80% threshold change about evidence collection?

It converts frequency from context into evidence. The T-TESS 2 Pilot Rubric, now posted publicly by Teach for Texas, states the rule directly: except where otherwise noted, for a performance level to be considered met, the descriptor must be met by the teacher or student at least 80% of the time during the observation or summative period. Where collecting evidence from every student is unrealistic, a representative sample of the class must meet the descriptor.

Under the current rubric, one strong instance of a practice could carry a rating, and often did. Under an explicit 80% floor it cannot. An appraiser who wrote "the teacher checks for understanding throughout the lesson" has recorded an impression that the new rubric has no way to score.

Several indicators go further and name the percentages outright. In D2.C Monitoring and Responding to Student Learning, the Student Progress indicator ties performance levels to the share of students whose responses or work show expected progress toward grade-level standards: less than 20% at Improvement Needed, 21 to 60% at Developing, 61 to 80% at Proficient, 81 to 95% at Accomplished, and 96 to 100% at Distinguished. That is not a judgment call your team can norm its way around. It is a counting task, and your team has to agree on what gets counted.

The rubric also layers upward. Higher performance levels are written as "AND" statements that add a new skill on top of the level below, and many of them shift from what the teacher did to what students did. An appraiser who never captured student actions in enough detail cannot support Accomplished or Distinguished under T-TESS 2, no matter how strong the instruction looked.

Can you provide some examples of T-TESS observation evidence under the new threshold?

The distinction every calibration session teaches is evidence versus inference: inference is what you concluded, evidence is what a camera would have captured. T-TESS 2 adds a third column. Evidence now has to carry frequency.

  1. Inference: "The teacher checks for understanding consistently." Old-style evidence: "Teacher used a cold call after the mini-lesson." 80%-threshold evidence: "Teacher checked for understanding at 5 of 6 transitions; skipped the transition into independent practice."
  2. Inference: "Students were engaged in the cognitive work." Old-style evidence: "Most students were working during practice." 80%-threshold evidence: "19 of 22 students were writing during the 12-minute practice block, 86%; 3 students at the back table had blank pages after minute 4."
  3. Inference: "Transitions were efficient." Old-style evidence: "Transition to stations was quick." 80%-threshold evidence: "4 of 5 transitions ran under 60 seconds. The transition to stations took 3 minutes and required three repeated directions."
  4. Inference: "The teacher used precise academic language." Old-style evidence: "Teacher modeled the term 'numerator' correctly." 80%-threshold evidence: "Teacher used precise academic vocabulary in 11 of 12 explanations; one substituted 'top number' for numerator during guided practice."

The third column is more tedious to write and it is the only version that survives a rating challenge under the new rubric. Our guide to writing better observation notes covers how to capture counts in real time without watching the tally instead of the lesson. The short version: pick two or three indicators before you walk in, and count only those.

What does a T-TESS 2 recalibration session look like?

Run the 90-minute norming protocol you already have, with one addition that makes it a recalibration rather than another norming session.

  1. Score the same recorded lesson twice. Once on the current 16-dimension rubric, once on the T-TESS 2 dimension your team is norming. Same lesson, same appraisers, two scoring passes.
  2. Examine every case where the rating moved. A rating that fell from Accomplished to Proficient under the new indicators is the most useful artifact your team will produce this year. Ask which specific indicator language caused the drop, and whether every appraiser saw the same cause.
  3. Sort the movement into two buckets. Ratings that moved because the rubric genuinely raised the bar, and ratings that moved because two appraisers read a new indicator differently. The first is information to carry into teacher communication. The second is the calibration problem, and it is the only bucket worth the rest of the meeting.
  4. Write down the interpretation. One sentence per resolved disagreement, naming the evidence that places that indicator at that level. Without the document, everyone leaves agreeing and returns in October back where they started.
  5. Practice counting before you practice rating. Spend one session where nobody assigns a level. Appraisers watch a lesson and produce only frequency counts on two indicators, then compare the counts. Teams are usually surprised by how far apart the raw counts are, which is a cheaper thing to discover in a training room than in a grievance meeting.

Double-scoring is the step that does the work. It surfaces exactly where your team's old shared standard fails to map onto the new language, which is information nothing in a recertification course will give you.

What should you norm on before your district adopts T-TESS 2?

More than you could a month ago. TEA has now posted the full pilot rubric, so the dimension labels, guiding questions, and target indicators are readable today: D1.A Lesson Design, D1.B Instructional Internalization, D1.C Supporting All Learners, D2.A Instructional Alignment, D2.B Instructional Content, D2.C Monitoring and Responding to Student Learning, D3.A Efficient Routines, D3.B Student Persistence, D4.A Professional Behavior, D4.B Goal Setting, Self-Reflection, and Professional Development, and D4.C Communication and Partnership.

Two cautions on that list. TEA's administrator correspondence describes 12 dimensions and the posted pilot document lays out 11, so count from the PDF rather than from any summary, this one included. And D1.A applies only to teachers not using SBOE-approved high-quality instructional materials, which means the dimension set your appraisers work from depends on your materials adoption.

The pilot rubric is also a pilot. It is being tested by 15 school systems during 2026-27 specifically so TEA can revise it. Norm on it, but do not print it on laminated cards.

How is recertification training different from recalibration?

Recertification training is TEA's. Recalibration is yours.

Recertification teaches appraisers what the rubric says and certifies that they can apply it. It begins in spring and summer 2027. Recalibration determines whether your specific team, watching the same lesson, lands on the same level for the same reason. TEA does not run it, fund it, or schedule it, and it is the line item that disappears when a professional development budget gets trimmed in March.

The practical consequence is a calendar problem. If your district adopts in 2027-28, appraisers finish recertification in summer 2027 and start rating in September. The only window for a recalibration session is the same August pre-service block that everything else competes for. Put it on the calendar this fall, while 2027 dates are still empty. For district leaders trying to see whether interpretation actually holds across ten campuses rather than one, district partnerships surface that pattern in the write-ups themselves.

Observation Copilot supports the current T-TESS rubric natively, organizing raw notes by domain and dimension and mapping evidence to the five-point scale. It does not assign ratings and it does not replace appraiser judgment. What it does is make every appraiser's write-up structurally comparable, so an interpretation gap between two administrators shows up on the page instead of hiding in two differently organized documents. Principals across Texas use it to get from notes to a structured draft the same day.

Observation Copilot has helped me to streamline and speed up the teacher feedback process. I'm able to organize the notes that I take during the lesson into summaries about the different dimensions while I get straight to the work of rating the lesson on the rubric.

- Jason Cunningham, Principal, Stockdale Independent School District, Stockdale, TX

That division of labor is the whole argument for doing calibration work before a rubric transition rather than during it. When the structure and the vocabulary are handled, appraiser attention stays on the judgment, which is the part that has to be relearned when the rubric changes. If you want models of what the finished write-up should look like, our observation feedback examples show the level of specificity the new indicators expect.

Frequently Asked Questions

What is a good T-TESS score under T-TESS 2?

The scale is unchanged: Improvement Needed, Developing, Proficient, Accomplished, and Distinguished, with Proficient representing solid expected practice. What changed is the threshold. A performance level descriptor must generally be met at least 80% of the time during the observation or summative period for that level to count as met.

How do you calculate a T-TESS score under the new rubric?

Each dimension is rated on the five-point scale from the target indicators beneath it, and higher levels layer additional requirements onto lower ones. Several indicators specify percentage bands directly. In D2.C Student Progress, for example, Proficient requires 61 to 80% of students showing expected progress and Accomplished requires 81 to 95%.

What are some good questions for T-TESS coaching conversations?

Ask what students did, not what the teacher planned. How many students demonstrated the objective, and how do you know? Which check for understanding changed what you did next? Where in the lesson did the fewest students engage? Under T-TESS 2 those questions are also rating-relevant, since higher performance levels are written in terms of student actions.

When does T-TESS 2 appraiser recertification training happen?

Spring and summer 2027, per TEA, with additional detail promised in fall 2026. Recertification is separate from calibration. Training certifies that an appraiser can apply the rubric; only a norming session run by your own team establishes that your appraisers apply it the same way as each other.

Can Observation Copilot help keep appraisers consistent under a new rubric?

Indirectly. It does not assign ratings or replace evaluator judgment. It organizes each appraiser's notes into feedback aligned to the same framework domains and indicators, which makes differences in interpretation visible rather than buried in inconsistent write-ups. It is free for individual principals at app.observationcopilot.com.

Give every appraiser on your team the same starting point for T-TESS 2.