Tuesday, January 10, 2012

Measuring teacher quality - adjusted "value addition"

Nobody disputes that teacher quality is central to improving student learning outcomes. But there is limited agreement on how teacher quality should be measured. The simplest and commonplace method has been to use standardized test scores to assess teacher quality. But its flaws are all too obvious - grade inflation, does not control for socio-economic backgrounds of children etc.

In this context, there has been considerable interest and debate surrounding the use of teacher "value addition" (VA) - by using their students’ test score gains - as the preferred metric to measure teacher quality. A teacher's VA is defined as average test-score gain for his or her students, adjusted for differences across classrooms in student characteristics (such as their previous scores). However, its critics have raised the same questions arguing that students are not randmoly assigned to teachers and therefore VA measures have the same flaws as standardized test scores.

A newly released longitudinal research study that tracked 2.5 million students in the US over 20 years by Raj Chetty, John N. Friedman, and Jonah E. Rockoff (presentation slide here) and assessed the value addition of teachers over the long term. Controlling for numerous factors, including student socio-economic backgrounds, they found that value-added scores consistently identified some teachers as better than others. They write,

"Are teachers’ impacts on students’ test scores ("value-added") a good measure of their quality? This question has sparked debate largely because of disagreement about (1) whether value-added (VA) provides unbiased estimates of teachers’ impacts on student achievement and (2) whether high-VA teachers improve students’ long-term outcomes. We address these two issues by analyzing school district data from grades 3-8 for 2.5 million children linked to tax records on parent characteristics and adult outcomes. We find no evidence of bias in VA estimates using previously unobserved parent characteristics and a quasi-experimental research design based on changes in teaching staff. Students assigned to high-VA teachers are more likely to attend college, attend higher-ranked colleges, earn higher salaries, live in higher SES neighborhoods, and save more for retirement. They are also less likely to have children as teenagers. Teachers have large impacts in all grades from 4 to 8. On average, a one standard deviation improvment in teacher VA in a single grade raises earnings by about 1% at age 28. Replacing a teacher whose VA is in the bottom 5% with an average teacher would increase students’ lifetime income by more than $250,000 for the average classroom in our sample. We conclude that good teachers create substantial economic value and that test score impacts are helpful in identifying such teachers."


Translated into English, as NYT writes, it appears staggering,

"Replacing a poor teacher with an average one would raise a single classroom’s lifetime earnings by about $266,000, the economists estimate. Multiply that by a career’s worth of classrooms. If you leave a low value-added teacher in your school for 10 years, rather than replacing him with an average teacher, you are hypothetically talking about $2.5 million in lost income."




As the graphic below shows, when a high value-added (top 5%) teacher enters a school, end-of-school-year test scores in the grade he or she teaches rise immediately.



A few observations

1. There cannot be any denying the fact that quantitative measures of student performance have to be at the center of any meaningful attempt to measure student learning outcomes and bring in accountability in teachers. Data analytics that mine this database can help draw decision-support inferences.

2. Even with all its flaws of cheating in tests, teachers cherry-picking students, and some good teachers taking the fall, VA measure may be, atleast for now, least worse among all bad frameworks to measure teacher quality.

3. The effectiveness and credibility of teacher VA measures can be improved by adding a qualitative dimension to assessment. Such assessments can take care of the several intangible factors that while critical to moulding students are not often captured in test scores.

4. Such assessment measures should not be used for high-stakes decisions. It should be used to identify deficiencies and put in place mechanisms to address them, through, for example, more focussed and targeted trainings for teachers etc.

5. Counter-intuitively, teacher VA measures may be more reliable and credible measure of teacher quality in government schools in countries like India because the socio-economic backgrounds of children are more or less similar. It is also much easier to capture major differences in social categories - religion, caste etc - and control for them in the VA measures.

6. Finally, such measures become more accurate as it accumulates years of data and larger sample sizes, which enable drawing more credible longer-term conclusions. In this context, it becomes possible to kick-in high-stakes decisions like say salary increments to senior teachers.

Update 1 (17/1/2012)

NYT Room for Debate on measuring teacher effectiveness using student examination results.

Update 2 (28/2/2012)

The New York City makes public value added ratings of 18,000 city school teachers amidst strong opposition from unions who lost the court battle to stop its release. The ratings, known as teacher data reports, covered three school years ending in 2010, and are intended to show how much value individual teachers add by measuring how much their students’ test scores exceeded or fell short of expectations based on demographics and prior performance. In 2010, The Los Angeles Times had hired a statistician and published its own set of ratings.

Such ratings have been gaining currency, in part because they are favored by the Obama administration’s Race to the Top initiative, which makes adoption of such measures a precondition for receipt of federal funds. New York City principals have made them a part of tenure decisions. Houston gave bonuses based in part on value-added measures, though that program was reorganized. In Washington, poorly rated teachers have lost their jobs.

Sunday, December 25, 2011

On trade numbers and statistical illusions

The Economist has an interesting debate on whether persistent trade deficits are a bad thing. Hal Varian makes an important point about the fallacy of paying too much importance to headline trade deficit figures.

According to research by Ken Kraemer at UC Irvine, the component parts of the iPad are imported to China from South Korea, Japan, Taiwan, the European Union, the US and other places for final assembly. None of the component parts are made in China: it's only role is assembly. The value added by the final assembly in China is about $10. Nevertheless, each iPad exported from China to the US increases the US trade deficit with China by $275.

The same misleading accounting holds for other products. If China buys steel, aluminum, and machine tools from Australia and uses these parts to build a ship which they then export to the US, the total value of the ship is counted as an export for China.


Laurence Kotlikoff and Scott Sumner argue that the most important metric should be the national savings rate, since it determines not only the sustainability of a trade deficit but also whether the deficit is financing productive investments.

On the same subject of statistical illusions, Paul Krugman posts that Ireland's reported recovery in competitiveness may not be a reflection of the true story.

Ireland is an economy that generates a lot of GDP — but not much GNP — out of capital-intensive, foreign-owned export sectors, such as pharma. And what has happened in the austerity era is that these sectors, which aren’t selling to the domestic market, have held up much better than labor-intensive sectors serving that domestic market. And this causes a spurious increase in labor productivity: if you lay off a construction worker but don’t lay off a pharma worker who basically watches over very expensive machines that produce a lot of output, it looks as if productivity has gone up, but in any individual sector nothing has happened.


In other words, the numerator (GDP) falls by a far smaller number than the denominator (workers) when a less productive domestic worker is displaced.

Update 1 (26/1/2012)

The Economist points to a study about iPad's production supply chain and writes about how trade statistics overstate trade figures

"According to a study by the Personal Computing Industry Centre, each iPad sold in America adds $275, the total production cost, to America’s trade deficit with China, yet the value of the actual work performed in China accounts for only $10. Using these numbers, The Economist estimates that iPads accounted for around $4 billion of America’s reported trade deficit with China in 2011; but if China’s exports were measured on a value-added basis, the deficit was only $150m."




China’s small contribution to total costs suggests that a yuan appreciation would have little impact on its exports. A 20% rise in the yuan would add less than 1% to the import price of an iPad. For imports such as clothing and toys the Chinese value added is much higher. But electrical machinery and equipment, with more complex cross-border supply chains, make up one-quarter of China’s exports to America.

Tuesday, November 29, 2011

Fiscal Policy Matters - Big Time?

Christina Romer is anguished about the fiscal policy debate,

"Policymakers and far too many economists seem to be arguing from ideology rather than evidence... the evidence is stronger than it has ever been that fiscal policy matters — that fiscal stimulus helps the economy add jobs, and that reducing the budget deficit lowers growth at least in the near term. And yet, this evidence does not seem to be getting through to the legislative process. That is unacceptable. We are never going to solve our problems if we can’t agree at least on the facts. Evidence-based policymaking is essential if we are ever going to triumph over this recession and deal with our long-run budget problems."


She writes about the inherent difficulty of estimating the impact of fiscal policy on people's consumption decisions because there are other things happening in the economy and other unanticipated factors which might influence the policy. She points to the omitted variable bias while illustrating the case with the early 2008 Bush tax cut. The contribution of this tax rebate to holding up household consumption spending was offset by the tumbling houseprices and resultant lowering of household wealth which in turn squeezed disposable incomes. In other words, the omitted variable bias skews our understanding of important relationships, nowhere more so than in our assessment of fiscal policy.

Prof Romer also points to tax cuts made when the economy is slipping into recession and tax increases in response to increased spending needs (like, say, war). In the former, unless tax cut was very large and for a sustained period, the outcome would still be weak growth - the tax cut would have only mitigated the adversity. In the later case, there is a behavioural sleight of hand - the impression gains ground that the tax increase caused increase in spending, whereas it was only a small contributor to an already rising spending.

In order to address the omitted variable bias, she and husband David Romer excluded from their empirical analysis the tax changes taken in response to economic conditions and confined their dataset to tax changes made for ideological reasons (Reagan tax cut, Clinton tax increase etc). The graphic below shows two estimates of the impact of a tax cut of 1% of GDP on real output. The red line shows the result using the conventional measure of tax changes — the change in cyclically-adjusted revenues. The blue line shows the estimates based only on the relatively exogenous tax changes we identified from the narrative analysis.



The graphic shows that limiting omitted variable bias results in larger and more statistically significant estimated impacts of tax changes. Prof Romer also argues that nobody has done a similar study (one that controls for motivation) to assess the true impact (on output) of a government spending increase. In its absence, there is no way to derive more definitive conclusions about the real impact of direct government spending increases nor adjudicate beween the relative superiority of tax cuts and spending increases.

She highlights the counterfactual problem of what would have happened without the ARRA in the US,

"The metaphor I find helpful is to a patient who has been in a terrible accident and has massive internal bleeding. After life-saving surgery to stop the bleeding, the patient is likely to still feel pretty awful and will have a long way to go before he is fully healed. But that doesn’t mean the surgery didn’t work. You have to judge the effect of the surgery relative to what otherwise would have happened. Without surgery, the patient would have died."


She points to three studies which, albeit incomplete, point to the ARRA playing an important role in propping up demand and ensuring that the economy did not get any worse.

She also points to similar omitted variable bias in the study of Alberto Alesina and Silvia Ardagna which has become the touchstone for advocates of expansionary contraction to argue that fiscal austerity is generally expansionary. They mined budget data for a large number of advanced countries over the past 35 years and identified large fiscal consolidations by looking for times when the cyclically-adjusted budget deficit fell sharply. They find that output tended to rise on average after these consolidations, especially those focused on reductions in government spending. She writes about the omitted variable bias in their analysis,

"Some of their fiscal consolidations weren’t deliberate attempts to get the deficit down at all. Rather, they were times when the budget deficit fell because stock price booms were pushing up tax revenues. Stock prices were a big omitted variable. They were driving the deficit reduction and were likely correlated with rapid output growth. This omitted variable made it look as though deficit reduction was expansionary, when it really wasn’t."


An IMF working paper, about which I blogged earlier, sought to address the deficiencies of the Alesina-Ardagna study and found that austerity programs hurt. Their dataset was limited to deliberate fiscal consolidation - for reasons unrelated to short-run macroeconomic developments - moves in 15 advanced countries over the last 30 years. They find that unemployment typically rose and output fell following such austerity programs.

Monday, October 10, 2011

Analyzing India's Cricket Debacle - A Black Swan Event?

I have been thinking of posting this all through India's disastrous recent cricket tour of England. It was Chris Dillow's excellent post about cognitive biases in football that finally got me around to writing it.

The dismal performance of India's cricketers has been variously attributed to IPL, England emergence as the successor to the great West Indian and Australian teams of the past forty years, India's "club-side" like bowling attack and the inability of its batsmen to cope with the swinging ball, and so on.

Without going into the merits of each of these, if we view this performance in its true perspective - the sheer magnitude of the defeat, the recent relative performances of both teams, and an individual assessment of the players from both sides who played in the series - none of the aforementioned explanations appear convincing.

Consider these. India's 4-0 defeat, apart from being its worst ever against England in 15 series there, was also its worst test loss margin ever. In fact, even the great West Indies, with all its great bowlers and batsmen, or Australia of the last two decades, could not inflict a test defeat of such magnitude, even in series with more tests. Undoubtedly much weaker Indian teams, both in batting and bowling, have performed more creditably against far better teams than the current English team, even in conditions atleast as adverse as that in the recent series. A logical performance-based explanation would lead us someway down the conclusion that the current English team is among best ever cricket team or conversely this Indian team is among the worst ever team assembled by the country!

In the build-up to the series, both teams had equally impressive recent test records. If anything, India's performance was superior, both in terms of the fact that its successes were for a longer period of time and was against slightly better opposition. The No 1 test ranking was a just reflection of India's superiority. Apart from its big success in Australia last winter, England's recent victories have been against the lesser teams (nothing in Pakistan, Sri Lanka, India, and South Africa).

Interestingly, it needs to be borne in mind that the same set of English bowlers have played in all the last three test series between the two countries, two of which were in England, and two of which were won by India and one was drawn. James Anderson (in four) and Stuart Broad (in two) led the English bowling attack on these tours. This brings us to the English bowlers themselves. While Anderson is arguably one of the finest exponents of swing bowling in friendly conditions today, his place among the greats of swing bowling is questionable. Stuart Broad's place was itself under threat, though it can be argued that his best years may be ahead. Take out the performance of the last one year, and the averages speak for themselves.

Man to man, given the fact that Graeme Swann was hardly a factor in the first three tests, the South African attack of Dale Steyn and Morne Morkel, against whom the same Indian batting line-up fared with great distinction less than a year back, is far superior. In terms of every imaginable measure of a bowler's art, Dale Steyn is far superior to James Anderson. Morne Morkel is similarly superior to Stuart Broad. Although England's third seamer, Chris Tremlett or Tim Bresnan or Steve Finn, is superior to South Africa's, the added presence of Jacques Kallis evens up things on this front. So, if South Africa's bowlers are superior to the English bowlers, there is something amiss about attributing the extraordinary English performance to the excellence of their bowlers.

I have three explanations for the triumphalism of English cricket writers,

1. Statistical coincidence - As Chris Dillow writes, events occasionally turn out such that one team enjoys the rare confluence of all fortunate factors, while the other team suffers the exact opposite. England had all its players playing at the peak of their form and free from injuries (and given their otherwise normal averages, it cannot be denied that England enjoyed one of the rare runs of all players being in great form), conditions favorable to its bowlers, its batsmen facing a weak and demoralized bowling attack, and so on. India had exactly the opposite - the injury toll, even with IPL workload, and the near complete batting failure, being inexplicable.

And once, the coincidence of factors align in such comprehensive manner, and one team starts to suffer, it is more likely that its confidence will deplete just as fast as that of the other will rise. A self-fulfilling spiral is triggered off. A statistical outlier will then get mistaken for something else.

2. We live in the present - The stellar performance of the same English bowlers, who not far back were whacked by all and sundry in the cricket World Cup (admittedly there is difference between test and one-dayers, but not so much as to merit such differential - Morkel and Steyn hardly suffered such consistent punishment even in one-dayers), raises the issue of how we should assess modern day cricketers.

The amount of cricket modern cricketers play means that their careers are more likely to be short and spliced with injury or fatigue interruptions. In the circumstances, there is a strong case for assessing players based on their current form. Given the competition and standards, even lesser mortals (say, Time Bresnan or Chris Tremlett) playing cricket today are likely to have a streak of great form for some period of time. Then law of averages catch up and they fall back to their mean career trajectory. Only the great players, and there appear to be only a handful of true greats playing now, can sustain their high-level performance for years.

3. Availability Bias - The 3-1 victory over Australia in 2009-10 was easily the greatest cricketing achievement for England in nearly four decades, if not more. It was preceded and succeeded by consistent performances by the team, albeit, as aforementioned, against the weaker teams. The whitewash of the World No 1 Indian team on top of all this naturally reinforces the positive feeling about the team and therefore the impression of an all-conquering team.

In fact, this is classic availability bias, wherein the immediate recollections of the team's performance gets disproportionate importance in the overall assessment of the team. The immediacy of the string of these successes, amplified manifold by modern media coverage, gave rise to the impression of an all-conquering team.

None of this is to denigrate the English achievement nor condone the pathetic performance of India. It is only to fill in a sense of perspective to the emotion charged reporting that dominated the richly deserved triumph of the British team.

Friday, July 1, 2011

India and the World

Superb graphics from If It Were My Home, representing cross-country comparison on certain social and economic indicators.

For example, the average American spends 78.1 times more on health care, and consumes 27.6 times more oil and 25.8 times more electricity than the average Indian. In comparison to the average Indian, the average Chinese spends 2.5 times more money on health care, consumes 5.3 times more electricity and 2.6 times more oil, and has 59.81 times more chance of being employed, and 66.4% less chance of dying at infancy.

The only flattering comparison for an Indian would be when made with this country!

Sunday, May 15, 2011

Westward shift in US population center of gravity

Interactive graphics can convey complex issues in a cognitively striking manner. Economix points to a US Census Bureau interactive map showing the shifting center of the country's population - defined as "the place where an imaginary, flat, weightless and rigid map of the United States would balance perfectly if all residents were of identical weight".



(click this for full image)

Like in the US, India's decennial census figures too are out now. Since census figures are a wealth of statistical information, its utility lies in the ease with which potential users can extract (it baffles me as to why GoI Departments do not provide for downloading their data in atleast Excel format) and render this data. Compare the respective efforts of the Census Commissionerates from the United States and India in this regard. The Census Commissionerate could take a leaf out of this when presenting its final 2011 census data.