Tuesday, September 21, 2010

A Primer on the Nashville Incentive Pay Experiment, Part 4

Part 4: What Can We Learn?


Previous posts:
Part 1: Background Info
Part 2: What to Look For
Part 3: Why it Matters


Despite what Rick Hess says, we won't learn "nothing" from the results of the study (but do read his post, as most of his points are good ones).  So what, exactly, will we learn from the results of the study?  It's hard to say exactly, but here are some things that we can and cannot learn from the study:

We can learn how individual teachers respond to a financial incentive offered for individual results.
We cannot learn how individual teachers, groups of teachers, or entire schools respond to financial or other types of incentives offered to groups of teachers or schools.

We can learn whether teachers will change their teaching in ways that will raise student test scores in response to an individual financial incentive.
We cannot learn whether teachers will change their teaching in ways that will increase student engagement, critical thinking, creativity or a myriad of other factors in response to an individual financial incentive.

We can learn whether middle school math teachers in Nashville were more likely to switch schools/districts or leave the profession if they were in the control or treatment group.
We cannot learn whether talented people across the country are more likely to become teachers, and subsequently remain in teaching, if there are performance bonuses in place.

We can learn whether, under this particular system, teachers who are "better" as measured by standardized tests across all three years tend to be rewarded on a year-to-year basis.
We cannot learn whether, under performance pay systems, better teachers (as measured in any number of ways) tend to be paid more.

In short, yes, there are rather severe limitations on what we can learn from this one study.  But, at the same time, I'd argue that we can learn more from this study than from most others.  Despite what Rick Hess, we will not learn "nothing of value," in part because different people value different things.  Hess might think that the teacher recruitment/retention aspect of performance pay is the most important, but plenty of others think the incentivizing of effort is the most important.

What does this mean for the interested observer watching from afar?  It means that your ears should perk up if when you hear strongly worded statements from both sides of the debate.  This study is one piece in the puzzle -- and an important piece, at that.  It's neither the end all and be all of research into performance pay nor an utterly useless waste of time that fails to inform the debate in the least.

A Primer on the Nashville Incentive Pay Experiment, Part 3

Part 3: Why it Matters

Previous posts:
Part 1: Background Info
Part 2: What to Look For


You may have noticed that I'm devoting a fair amount of attention to the results of the Nashville incentive pay experiment that are being released today.  Let me take a couple of minutes to explain why.

The first, and most obvious, point is that this is the first randomized field trial evaluating the effectiveness of a merit pay system.  The debates to date on whether or not we should use some form of performance pay in school have largely relied on ideology and theory.  This will give us the first concrete, empirical, and comprehensive evidence to inform our future policy decisions.  Given the importance of merit pay in the national discussion right now, it makes this one of the most important education studies of the decade.

Now, that is not to insinuate that there still won't be a ton of unanswered questions about merit pay after the results of this are digested (no one study is ever enough to close the book on such a wide-ranging topic) but, rather, that we will know significantly more about how merit pay plays out at the ground level after the release of this study.

Or, at least, we certainly hope we will -- especially given that this represents countless hours of effort by dozens of people over the past five years or so . . . and millions of dollars.  Things can always be done bigger and better, but there won't be anything bigger and better than this for quite some time (if ever), so expect the results to be bandied about by both sides of the debate for years to come.  In other words, expect this study to be the definitive study on how individual teachers respond to financial incentives well into the future.

My next post will address what, exactly, we might learn from the results.  It will likely be followed up at an attempt to live-blog the release of the results beginning around 12:30pm central time.

Monday, September 20, 2010

A Primer on the Nashville Incentive Pay Experiment, Part 2

Part 2: What to Look For

The results from the Nashville incentive pay experiment are due to be released tomorrow (see last week's post for background info on the experiment) -- here are a few things to keep an eye out for in the final report:


Stability of scores

One of the issues with value-added scores has been their high variability from year to year.  Researchers were worried about the effects of "statistical noise" and random variations in scores before the start of the experiment.  In practical terms, you'll want to know how many teachers received a bonus each year versus how many earned a bonus one or two out of the three years -- and how well teachers who ever earned a bonus did across all three years (e.g. did they earn a bonus for being in the 85th percentile one year but then have a score in the 40th percentile the other two years).  In other words, were bonuses going to the same teachers who tended to outperform other teachers, or were they just randomly assigned each year?


What, if anything, did teachers do differently?

To me, this is the most interesting question.  Did treatment group teachers report putting more effort into teaching because of the incentive?  Did they assign more homework?  Did they crack down more on misbehaving students?  Did they work harder to get certain students out of their class?  Did they focus more on their math class than their other class (some of the teachers taught multiple subjects)?  Did they attend more PD?  Did they spend more time on test prep?  Or did they not bother to change at all?  The baseline numbers are going to get all the attention (i.e. did students of treatment group teachers score higher?) but, I think answering this question is far more important -- both because it informs us as to why, or why not, treatment teachers do better and because it gives us valuable insights into how teachers think and how their actions influence student achievement.


Teacher reaction to receipt/non-receipt of bonus

Did teachers who earned a bonus enthusiastically redouble their efforts while those who didn't simply gave up and maybe quit?  Or did those who failed to earn a bonus decide to redouble their efforts to make sure they got one the next year while those who earned one became cocky and put their teaching on cruise control?  In terms of the numbers, it will be interesting to see if there's any divergence between the scores of first-year winners versus first-year losers (in terms of bonus receipt) -- does one group go up in second years and other down, are both steady, or something else?  Of course, if the scores are simply random each year, then we'd expect first-year losers to do better than first-year winners in the 2nd/3rd years because of regression to the mean.


Demographics of winning/losing teachers

Bonuses are not based on a value-added measure that attempts to control for everything and isolate the individual teacher's effect.  The only thing the computed score takes into account is prior achievement of the students.  So it will be interesting to see if winning teachers taught in better schools, were more experienced, had smaller classes, used a different curriculum, had lower student transience, or other factors that might influence students' gains on the state math test.


Performance in different types of math classes

Were teachers of advanced algebra classes more or less likely to earn bonuses than teachers of remedial math classes?  Since the formula for determining bonuses is somewhat simplified, it may be the case that it's a lot easier to get low-performing student to advance x points than high-performing students, especially if there's a ceiling effect.  Or it may be the case that some subjects are better aligned with the contents of that year's state test.


Improvement of scores

For merit pay systems to transform our schools, we need teachers to improve -- and continue to improve -- while they're eligible for these bonuses.  Why?  Let's say that the treatment teachers, as a group, have a score that ranks them in the 50th percentile historically each year.  Now let's say that the treatment group teachers average a score in the 60th percentile each year.  That would appear to be strong evidence that incentives make teachers better (at least as measured by gains on Tennessee's state tests, anyway), but it would only offer a limited amount of hope for the future because it would indicate only a one-time boost to scores.  Let's say another experiment on a professional development system of some sort yields growth instead of simply a step up -- teachers are in the 50th percentile the first year, the 55th the second, the 60th the third, and the 65th the fourth . . . that system would offer more promise for future growth in school success.


Keep your eyes open for some more pre- and post-release analysis of the experiment . . .

Friday, September 17, 2010

Today's Random Thoughts

-The results from the Nashville incentive pay experiment are being released on Tuesday, so keep your eyes open . . . and check back here for some more pre- and post-release thoughts.

-Hopefully everybody read this post entitled "Can Exercise Make Kids Smarter?".  The short answer: yes, it can. There's a fast-growing body of literature linking academic performance to all sorts of social conditions and environmental factors that I think we should start paying more attention to.  In the end, though, the bigger question will be what we do with this knowledge (e.g. figuring out how to get kids to exercise more is probably harder than figuring out how exercise affects brain development, body function, and academic performance).

-There seems to be an awful lot of gnashing of the teeth over Fenty's loss and how it proves that mayors that tackle education reform can't get re-elected.  Isn't it instead possible that the objections of the public were more over style than substance?  Maybe people like Rhee's reforms but don't like the way she implemented them (or maybe they would just rather she implemented different reforms).  I think we'd need a heck of a lot more analysis before we could conclude that Fenty lost simply because he tried to change schools -- and that any other mayor who tries will share a similar fate.

Wednesday, September 15, 2010

The Most Surprising Result of the TIME Poll

I'm somewhat skeptical of national polling on most educational issues, since the public doesn't truly understand much of what they're being asked about (heck, half the wonks don't really understand NCLB).  So I read a few posts about the TIME poll, bookmarked it, and forgot about it.  Finally got around to reading it today . . . there are mixture of results that anybody can use to support their pet reform or ideology, but I think they're all at least fairly close to what we'd expect.  Only one question really surprised me:

4. What do you think would improve student achievement the most?
More involved parents: 52%
More effective teachers: 24%
Student rewards: 6%
A longer school day: 6%
More time on test prep: 6%
No answer/don't know: 6%


Talk about flying in the face of recent rhetoric . . . if we took the last 100 editorials and op-eds on education policy from the nation's major newspapers, how many would call for better teachers and how many would call for more involved parents?  I can only remember one recent one that sort of called for the latter (Samuelson's op-ed about motviation), so I might have to guess 99 to 1.

A Primer on the Nashville Incentive Pay Experiment

Part 1: Background Information

According to Eduwonk, results from the Nashville incentive pay experiment are due to be released soon.  I've been meaning for a while now to write up some background information on the experiment so that we have some context when the results are released, so this seems like as good a time as any.

The National Center on Performance Incentives was started in 2006 with a 5 year, $10 million grant received from the Department of Education's Institute for Education Sciences.  The center is housed at Vanderbilt University's Peabody College and run in conjunction with various partners, including the RAND Corporation and the University of Missouri.  Peabody's Matthew Springer and James Guthrie (now of the George W. Bush Institute for Public Policy) are the directors, and the center is staffed by people from a range of institutions across the country (full list).  The funding was to cover two experiments plus other related costs.  The first experiment was conducted in Nashville from 2006-09 and was dubbed the Project on INcentives in Teaching (POINT).

The center started at Vanderbilt the same time that I did, and I worked there during my first year (2006-07) to earn my keep around here.  I haven't been involved with the center since then and have no information on what the results are.

The original experiment design was to encompass 200 middle school math teachers in the Metropolitan Nashville Public Schools -- 100 in the control group and 100 in the treatment group.  Teachers in the treatment group were eligible for bonuses of up to $15,000 for each of three consecutive school years.  Each teacher received $750 every year for participating as long as they completed all the required surveys, interviews, etc.  Teachers were recruited into the experiment in the fall of 2006, not long after the school year had begun.

Bonuses were based on student gain scores* (not quite the same as value-added, see technical note at end) on the Tennessee state test (TCAP).  Unlike virtually every state, TN's assement is system is vertically scaled, meaning that scores can be compared across years on the same scale (a score of, say, 250 in 7th grade means the same thing as a score of 250 in 6th grade).  This means that a student who goes from 240 to 260 from 6th to 7th grade gained 20 points.  Meanwhile, researchers looked at the years preceding the experiment to determine the average growth of students at each level.  Taking the previous example, let's assume that the average TN 6th grader scoring a 240 on the state test then scores 255 next year.  This would mean that a student who scored 260 was 5 points above average.  For that, a teacher would receive a score of +5, and each student the teacher taught would be scored similarly.  The average score for a student with teacher x would be calculated.  The purpose of calculating scores this way was to strike a balance between statistical rigor and transparency/ease of communication.  The result is a calculation that's not quite as rigorous as a value-added score, but a lot easier for teachers to understand.

When the teacher's final score has been calculated, it's then compared to the historic average for middle school math teachers in Nashville.  If a teacher scores in the 80th percentile, they earn a $5,000 bonus, the 85th percentile earns a $10,000 bonus, and the 95th percentile yields at $15,000 bonus.  The targets for the bonuses stay the same the entire three years, so it's possible for every teacher in the treatment group to earn a bonus each year (in other words, they're not competing against each other).  It's my understanding that for the first year the bonuses were distributed along with paychecks the following fall, but I don't know what the procedures were the following two years.

The experiment ended in May, 2009 and a large team of researchers have been poring over data from test scores, interviews, surveys, and other sources of information ever since.  This means that there is going to be a lot of analysis released at some point in time -- and that it's going to take a while for even the most informed reader to sort through.


technical note: A "gain score" is simply the gain in a student's score from year to year (260 - 240 = a gain of 20 points), while a "value-added score" is an attempt to isolate a teacher's effect on a student's score and might control not only for a student's previous achievement level but also the other teachers he/she has or has had, the school he/she attends, demographic factors, class size, peer effects, and any number of other things.  In other words, a gain score is just the raw growth a student exhibits while a value-added score is a more precise estimate of exactly how a specific teacher influenced that growth (though value-added could be computed for schools, states, etc. as well).

Tuesday, September 14, 2010

Today's Random Thoughts

-Jay Mathews echoes a point I've often made in private: most news stories about the cost of college attendance grossly overstate what the average student actually ends up paying.  Though student loan debt is not a trivial problem, I think there are probably more people scared off by misperceptions of the costs of college than there are people who are bankrupt because of their attendance.

-Aaron Pallas shoots holes in the claims made in a recent op-ed about the miracles worked by a group of CA schools.  In an op-ed I've seen mentioned numerous places, Caitlin Flanagan claims that the ICEF elementary schools closed the achievement gap.  Pallas does some number crunching and finds that's not even remotely true (except for one of the five schools, and only in 2nd grade reading) -- indeed, their students' test scores are only slightly better than district averages for African-American students.  I have to say I'm mildly surprised that her claims made it past the editor's desk.  There are a lot of reasons to be skeptical of the charter school movement writ large, but there's really no arguing the fact that some charters have achieved outstanding test results.  In other words, there's plenty of statistical evidence to support arguments for the proliferation of charter schools -- it seems odd that anybody would need to resort to misrepresenting the test scores of a few select charters.

-Stephen Sawchuck makes a reasonable point about the possibility that value-added scores can save the jobs of unfairly maligned good teachers as well as unfairly maligning good teachers.  Both sides of the debate would do well to remember that there are many positive and negative aspects of value-added scores.

-Kevin Carey writes about a very interesting chart on the growth in college expenditures.  Basically, the chart shows that "student-oriented" expenditures at the top 1% of most selective colleges have skyrocketed, have grown quite quickly at other schools among the top 10%, and haven't grown terribly fast at the rest.  I suppose the takeaway points are that the growing concern about runaway spending in higher education really only apply to a select few colleges (where money isn't really an issue in a lot of ways), and that our colleges are growing further and further apart in terms of resources.

Oh Meyer Goodness! Redux

Last week, Peter Meyer wrote a piece that I called "baffling".  Well, at least he's consistent, because today he did the same thing again.  Over at Flypaper, he almost seems to be calling for a return to segregated schools.

Here's some context: This NY Times article today referenced this report about suspensions in urban middle schools, the major finding of which was that black students (particularly males) were much more likely to be suspended than white students -- and that the gap had widened over the past few decades.

Meyer makes a leap at the end of his post, implying that the desegregation of schools is responsible for this disparity and referencing an MLK quote from decades ago to back up his support of segregated schools.

In the post where he first references the MLK quote, he does make a reasonable point that desegregation shouldn't be our sole policy aim (though, at the same time, I don't think very many people think it should be).

But there are two major problems with his latest post:

1.) As far as I can tell, the report says zero about any differences in suspension rates between more and less racially diverse schools.  In other words, there's really no readily apparent evidence for Meyer's claim.

2.) 1954 has come and gone.  We, as a society, have decided that separate but equal is inherently unequal.  Nostalgia for the past is one thing, but do we really have to go back and repeat all our mistakes?  Black males are also more likely to be arrested, does that mean we should create segregated neighborhoods to accompany our segregated schools?

I hardly think Mr. Meyer's next post is going to argue for separate drinking fountains and bathrooms, but he'd do well to remember that it's a slippery slope.  If you're going to advocate for segregated schools, please do so more thoughtfully and use actual evidence.


update: I originally misspelled Mr. Meyer's name as "Mayer" in this post. My apologies; no matter how much I disagree with him on this issue he still deserves to have his name spelled correctly.

Monday, September 13, 2010

Responsibility With No Responsibility

Researchers and practitioners all seem to agree that teachers are the most important factor within a school.  And many have taken that another step and asserted that teacher quality is almost the only thing that matters.  I've pushed back by pointing out that lots of things affect a teacher's performance other than a person's talent or moral character.

But here's what blows my mind.  Across the country, people seem to argue that teachers need to put up or shut up -- and that if a school fails it must be a result of poor teaching.  And, yet, all across the country, teachers are told to do things the way the principal, superintendent, board of ed, or whomever wants them done.  The last decade has seen the proliferation of scripted curricula ("teacher-proofing" they call it) and increasing micromanagement in urban schools (ask an NYC teacher if they have their "word wall" up or if their bulletin board properly displays student work).  If teachers bear all the responsibility for student success, why are they given so little responsibility for what and how students learn?

Think about it: if a teacher's not given any responsibility for how and what their students learn, then how can we hold them responsible for how and what students learn?  It's accountability without autonomy, responsibility with no actual responsibility.

When NYC started their principal accountability program, it was in the context of an "autonomy zone".  Principals signed contracts that basically said they would be fired if student achievement didn't improve in 5 years.  And, in return, principals had far more say over how their school was run and how professional development funds were spent.

Teachers, on other hand, aren't really offered the same deal.  They're essentially being told that they will be held responsible for what happens in their classroom (which isn't entirely unfair) -- but also that they will run their classroom a certain way . . . or else. 

If we don't trust teachers to do what's in the best interest of students, then maybe they're not the ones we should be pointing fingers at when students don't learn.  If a teacher follows a scripted curriculum and students don't learn, maybe we should point our fingers at the curriculum writers.  If a teacher follows the checklist the district passes down and students don't learn, maybe we should point our fingers at the district personnel.  If a teacher does everything their principal demands of them and students don't learn, maybe we should point our fingers at the principal.

If we think a teacher's primary responsibility should be to stick to the curriculum, decorate their rooms the way the superintendent says to, and follow the instructions of their principal, then, by all means, we should evaluate them on these things and hold them accountable when they fail to do them.  But if, instead, we think a teacher's primary responsibility is to ensure that students learn, maybe we should think about letting them determine what and how students learn before holding them accountable for this.

Friday, September 10, 2010

Today's Random Thoughts

-Tennessee has figured out a solution (hat tip: Stephen Lentz) to the fact that only about 1/3 of teachers teach tested subjects but that all teachers are supposed to have 35% of their evaluation based on value-added scores . . . all "non-TVAAS" (the state test) teachers will simply have their school's average score used for their evaluation.  Problem solved!

-Aaron Pallas continues his critique of the LA value-added kerfuffle, arguing that the LA Times did not do enough to inform its readers about the statistical uncertainty in value-added measurements.  He argues that they should've used confidence intervals (something that popped into my head the other day) to more accurately describe the estimate of a teacher's effect on student test scores (they send you a confidence interval with your SAT scores, so why not with a value-added score?) in addition to better describing year-to-year and subect-to-subject variability.  This is a follow-up to his incisive critique of the Times' failure to follow normal standards of journalism when verifying the student data.

Jay Mathews has Killian Betlach's take on what it's like to be told to restructure a school.

Roger Garfield, a teacher in DC, provides an insider's view of some of the problems the schools face.  The first couple paragraphs brought back a lot of memories for me.

Newark's answer to the Harlem Children's Zone is the Global Village, a group of five schools that have received federal turnaround dollars.

Robert Samuelson says the real key to reform is student motivation.  It's a pretty short op-ed, and there's a lot more to it, but I think he raises a valid point.  If student motivation doesn't change, why would we expect student learning to change?  But I don't think it's quite as strong of a repudiation of other policies as he argues, since better principals, better curricula, better teachers, smaller classes, and so on could conceivably alter student motivation (but if they don't, they probably won't work).