As I was perusing The New Bill James Historical Abstract I ran into the article "Bird Thou Never Wert" where James says:
"There is simply no such thing as a starting pitcher who has a long career with a low strikeout rate...I am always amazed that people fail to see this on their own. The career expectation for a strikeout pitcher is so much greater than the career expectation for a non-strikeout pitcher of the same age and ability that the difference is very obvious if you study the issue."
James then goes on to set the bar for a non-strikeout pitcher at roughly 4.00 strikeouts per nine innings.
Because the Royals have a young non-strikeout pitcher in Jimmy Gobble who in his career has struck out 3.50 batters per 9 innings (167 innings, 65 strikeouts), I wanted to see just what that difference was.
To do so I found a group of 69 non-active pitchers since 1945 that started more than 10 games, pitched more than 100 innings, were between the ages of 18 and 23, and had an ERA lower than the league average. This criteria gives pitchers of the same age and same ability at the beginning of their careers. Their cumulative stats were:
ERA AVG W WHIP ERA+ Career W Years Innnings K/9
3.21 9 1.27 84.7 92 11 1557 5.67
where AVG W is average wins in their qualifying season, Career W is their Average Career Wins, Years is the number of years pitched in the majors, Innings is their Average Career Innings, and K/9 is their stirkeout rate in their qualifying year.
When you break these down into pitchers who struck-out fewer batters than the league average versus those that struck out more (I didn't feel using 4 batters per game was appropriate since the league strikeout rates over time vary quite a bit) you get the following two groups.
# ERA AVG W WHIP ERA+ CareerW Year Innings K/9
30 3.25 8 1.28 84.6 56 8 1003 4.44
# ERA AVG W WHIP ERA+ CareerW Year Innings K/9
39 3.18 10 1.26 84.8 120 13 1983 6.61
The most striking thing about these two groups of course is that strikeout pitchers (the second group) won more than twice as many games in their careers as non-strikeout pitchers and had careers that were on average five years and almost 1,000 innings longer. And what's more, as you increase the distance from the league average the gap only gets larger (which James also notes). For example, when looking at only those pitchers +-.25 strikeouts per 9 innings the non-strikeout pitchers have career expectations of 54 wins and 7 years while the strikeout pitchers have 130 wins, 14 years, and over 2,100 innings pitched.
Certainly there a few exceptions in the list of 30 non-strikeout pitchers, the most notable being Joe Niekro who won 221 games over 25 years. Remove him from the list (which is arguably the right thing to do since he was a knuckleballer) and the non-strikeout group moves down to 51 wins and 7 years. The top two pitchers that remain are Bret Saberhagen with 167 wins and Scott Erickson with 140 wins.
So it appears that James was correct. Another way to look at this issue is to look at pitchers with long careers. I selected all those pitchers whose careers started after 1945 who've pitched more than 15 years in the big leagues. Interestingly there were exactly 200 such pitchers. Taken as a group these pitchers struck out an average of 5.9 batters per 9 innings in their first four years when the league average was 5.3.
Why is this the case? James speculates that there are two reasons. First, he says that generally speaking "the batting average against a pitcher is inversely related to his strikeout rate." This is the case since fewer balls are put into play for strikeout pitchers and as we've learned in the past few years with DIPS. Most pitchers don't have much control over the rate at which balls put into play turn into hits. This means that strikeout pitchers can afford to do other things less well (control for example or giving up homeruns) and still be successful. This gives strikeout pitchers more wiggle room as their physical skills decline as they get older. Once they dip much below the league average (historically between 4 and 6.75 from 1950-2004) they are on the way out. As an aside I should mention that the batting average on ball put in play (BABIP) for the sets of pitchers shown above was almost identical at .289 and .290. Secondly, a non-strikeout pitcher is "walking a fine line" since he doesn't have the same wiggle room as the strikeout pitcher. As hitters adjust or the pitchers lose a little ability they are hit harder and are then on their way out.
I don't disagree with either reason. A good general manager should look for young pitchers with decent strikeout rates and trade those who have to walk the fine line to be successful. As far as Jimmy Gobble is concerned, if the Royals bring him back up and he pitches decently I would trade him as soon as possible.
Friday, August 13, 2004
Gobble and K's
Posted by
Dan Agonistes
at
9:05 PM
0
comments
Thursday, August 12, 2004
MLB Pocket Manager
When should a major league manager attempt to steal? When should he attempt to bunt? How does that change in the late innings of a close game?
These are all questions that baseball fans have long thought and argued about. As Alan Schwarz shows in his excellent book The Numbers Game, the first real attempt at answering that question came in the late 1950s when father and son Charles and George Lindsey manually scored and tabulated over 1,000 games. From this data George went on to publish a 24-page article called "An Investigation of Strategies in Baseball" in the journal Operations Research.
In that article Lindsey included tables that show the number of expected runs and the probability of scoring in the remainder of a half-inning given the 24 base-out combinations and used them to answer the questions posed above. John Thorn and Pete Palmer produced an updated version of the table looking at data from 1961-1977 in their 1984 book The Hidden Game of Baseball and importantly also included a similar table of the probabilities of the home team winning in the bottom of the 7th down by a run. With the availability of play by play data starting in the mid 1980s other tables that show the probability of scoring one or more runs have been developed from play by play data a number of times and discussed, for example, in Curve Ball by Albert and Bennett show a table for 2002 data. One such set of tables calculated for 1999-2002 is provided by Tangotiger on his site. Another great computer-ready Win Expectancy table using data from 1979-1990 can be found here.
It always amazes me that this information is not more widely distributed and used in telecasts that regularly pack the screen with useless stats based on ridiculously small sample sizes (Matt Stairs is 3 for 11 when leading off an inning on Tuesdays) or that are simply meaningless (Corey Patterson has a hit in 24 of his last 30 games). So that I could watch a game and evaluate the odds of scoring and wining using various strategies I wrote a small .NET Compact Framework application for the Pocket PC called MLB Pocket Manager. This app uses the run potential, scoring probability, and win expectancy tables from Tangotiger. Using these tables it's simple to evaluate, for example, whether a manager should steal second with 1 out. The calculation goes like this:
Run Potential Before the Attempt = .573
Scoring Probability Before the Attempt = 28.3%
Run Potential If Successful = .725
Scoring Probability If Successful = 40.6%
Run Potential If Caught = .117
Scoring Probability If Caught = 7.7%
So the break-even percentage, the stolen base percentage that the runner needs to exceed, to increase the run potential and the scoring probability can be calculated as:
BErp = (.573 - .117) / (.725 - .117)
BEsp = (.283 - .077) / (.406 - .077)
The results: the break even percentage for increasing the run potential (BErp) is 75% while for increasing the probability of scoring (BEsp) it is 62.6%.
The same calculation can be made using the Win Expectancy table. For example, in the bottom of 7th with 1 out the win probability for the home team is 59%. The break-even can then be calculated as 65.7%.
In MLB Pocket Manager the strategies included are Steal 2nd, Steal 3rd, Double Steal, Sacrifice, Attempt an Advance on a Fielder's Choice, and Tag up to advance.
Here is the simple user interface for the situation described above.
The second tab includes the tables for quick reference.
The third tab describes the methodology.
You can download the CAB files for the various processor types here. Use at your own risk and you'll need the .NET Compact Framework v1.0 installed on your device first which you can get here.
Posted by
Dan Agonistes
at
11:45 AM
0
comments
Lost in Left Field
A fellow SABRite and fan in my neck of the woods shares his thoughts. Good stuff.
Posted by
Dan Agonistes
at
2:31 AM
0
comments
Tuesday, August 10, 2004
Percentage Plays
Just a quick note on tonight's 8-6 Cubs loss to the Padres. Besides the frustration of the Cubs once again hitting homeruns with the bases empty (Sosa, Lee (2), Garciapara, Alou) and Mark Prior again showing no control and getting squeezed by Bruce Froemming, the 9th inning included one of my pet peeves.
After Lee had homered to bring the game to 8-6, Barrett walked with one out and was there with 2 outs as Patterson came to the plate. Nevin at first base played well behind Barrett and yet Barrett did not take second. There are two good reasons why Barrett should have taken the free base:
- Taking second eliminates the force at second and with a speedy hitter like Patterson may allow the inning to continue on an infield hit
- Taking second clears second base in the event that Patterson does get a hit which allows him to steal second putting the tying run in scoring position
Of course scenario number 2 happened as Patterson singled up the middle. Barrett would likely have scored making it 8-7 with Patterson perched on first. Since Ramon Martinez was the pinch hitter it would have made even more sense for Patterson to steal since 2 hits were likely required to score. In fact, the breakeven percentage on scoring a run in that base-out situation is only 61%. Once again a major league manager fails to make the percentage play.
Posted by
Dan Agonistes
at
9:02 PM
0
comments
Royals Sabermetric Report
This is first of periodic sabermetric reports on the Royals and Cubs. Basically, this report is meant to fill in the statistical holes you'll find in newspapers and the general media. I've broken down the report into Team, Offense, Defense, Pitching, and Management that includes stats and a comments section based on sabermetric principles.
As of Monday August 9th
Team
Actual Record: 39-70 467 Runs Scored, 609 Allowed
Pythagorean Record: 41-69
Kauffman Park Factor: .933 for runs (18th) and .767 for homeruns (27th)
Comments: Last year the Royals outperformed their Pythagorean projection by 4 games and this year they've regressed towards the mean and then some. Only the Diamondbacks, Expos, and Mariners have scored fewer runs and only the Indians, Reds, Rockies, and Diamondbacks have allowed more. Poor offense + poor pitching = a 10-58 record when scoring 4 or fewer runs and 14-59 when allowing 4 runs or more. Last year they were 68-34 when scoring 4 or more runs and 31-76 when allowing 4 or more runs.
The combination of moving the fences back and the cooler, wetter weather in Kansas City this year has made Kauffman a pitcher's park again after several years being a hitter's paradise. I don't think the park changes have played any role in the Royals poor year. Moving the fences back should have helped May and Gobble which likely offset the damage it did to Sweeney, Randa, and Stairs.
Offense
OPS Leaders
Sweeney 850
Stinnet 836
Stairs 811
Harvey 807
Randa 721
VORP Numbers
Sweeney 24.8
Harvey 20.7
Stairs 12.5
Randa 5.1
Stinnet 4.7
Graffanino/Santiago 4.5
Berroa 1.2
Buck -6.4
Relaford -7.9
Brown -9.8
Pitches/Plate Appearance
Randa 3.89
DeJesus 3.88
Stairs 3.87
Harvey 3.76
Berroa 3.61
Brown 3.61
Relaford 3.72
Sweeney 3.40
Comments: Sweeney, despite having a bad year is still the Royals best offensive player in OPS and value over replacement player (VORP). Ken Harvey continues to be overrated with his .300 AVG still intact with only 33 extra base hits and 23 walks. Not suprisingly he also has the highest ground out to fly out ratio on the team at 2.00 and has grounded into a team high 12 double plays. The book on Harvey is to bust him low and inside with hard stuff. When opposing teams do that he hits plenty of foul balls in the batter's box and soft grounders in the infield.
Only Sweeney, Harvey, and Stairs actually have much positive offensive value. Dee Brown, despite only 143 plate appearances somehow managed to be -9.8 in VORP. This is due to his 5 (yes only 5) extra base hits and 4 walks for an OPS of 553. I was one who advocated giving Brown a chance to play this season to see what he could do. The Dee Brown experiment should mercifully end when the season does.
Sweeney is seeing fewer pitches this year than previously (over 3.70 the previous 3 years) , which accounts for his walks being down. He's also flying out more than in the past but I'm not sure what if anything that means.
As a team the Royals have a .318 OBP, last in the league. The patience that hitting coach Jeff Pentland preaches has not taken hold.
On the plus side Abraham Nunez is an exciting player to watch and should garner the lion's share of at bats the rest of the way to see if he'll figure into next year's plans. He still has a chance to be a productive major leaguer. I would install him in right field everyday and then share time between Brown and Mateo in LF (or bring Guiel back up and release Brown). In addition, David DeJesus has an 813 OPS since the break and has seen 3.89 pitches per plate appearance. He should only get better and should be fun to watch in the upcoming years.
Defense
Defensive Efficiency: .6787
Zone Rating
Graffanino 2b .804 (7th)
Berroa ss .768 (10th)
Randa 3b .813 (2nd)
DeJesus cf .846
Stairs rf .884
Sweeney 1b .833
Harvey 1b .818
Relaford 2b .745
Comments: The Royals are tied for last in efficiency (the odds of making an out when a ball, other than a homerun, is put into play) with Cleveland. Coming into the season Tony Pena thought the team had a good defense. Boy was he wrong. The problem with looking at Zone Rating (the odds of making a play on a ball hit into the fielder's zone) for the Royals players is that many of them haven't played enough innings to get a good sample size. However, Berroa, Relaford, Sweeney, and Harvey all have very poor zone ratings while DeJesus, Stairs(?), and Graffanino's are ok (remember that zone rating's differ based on position). Only Randa can be said to be a good defensive player on this team.
Pitching
VORP Numbers
Grienke 15.1
Cerda 13.3
Camp 9.4
Sullivan 9.4
May 7.1
Reyes 6.7
Field 6.7
Wood 5.7
Gobble 1.0
Anderson -18.2
WHIP
Greinke 1.147
Field 1.315
Cerda 1.324
Camp 1.341
Gobble 1.373
Wood 1.396
May 1.481
Reyes 1.547
Sullivan 1.555
Affeldt 1.663
Anderson 1.720
BABIP
Greinke .257
Cerda .260
Field .260
Gobble .269
May .308
Wood .318
Camp .328
Reyes .335
Anderson .335
Affeldt .350
Sullivan .352
Comments: Of the starters only Greinke is pitching well with Reyes, May and Wood holding their own but suprisingly the relievers (Cerda, Field, Camp, Sullivan) have all pitched relatively well. The average on balls in play might indicate that Anderson has indeed been the recipient of bad luck (which seems to be reversing itself these last 2 starts) while Gobble and Greinke have had their share of luck.
Management
SB: 54/85
SAC: 27 (tied for fewest in the AL)
IBB: 25 (10th in the AL)
Comments: When you don't count Beltran the Royals are 40 of 68 in stolen bases (59%) which has certainly cost them a few runs. Pena has been bunting more of late but fortunately they still haven't bunted often and have not walked many batters intentionally.
Calvin Pickering in Omaha (AAA) is second in the PCL in Runs Above Replacement Position (RARP) and has a Major League Equivalent Average of .289 (.260 is average). In fact, his MLEQA is higher than anybody on the Royals roster. Sweeney is at .282, Harvey and Stairs at .272, and Randa at .255. This may be the first time in history that a team has a player in the minors who has actually outperformed every player on their major league roster. Pickering has more than earned the right to be playing in Kansas City.
Posted by
Dan Agonistes
at
9:00 AM
0
comments
Productive Outs
This season ESPN and the Elias Sports Bureau have begun tracking "Productive Outs". These are defined as:
"A Productive Out, as defined and developed by ESPN The Magazine and the Elias Sports Bureau: when a fly ball, grounder or bunt advances a runner with nobody out; when a pitcher bunts to advance a runner with one out (maximizing the effectiveness of the pitcher's at-bat), or when a grounder or fly ball scores a run with one out. "
It has been suggested that tracking productive outs is part of a backlash against the Moneyball movement (emphasize on base percentage, don't throw away outs on sacrifice bunts and stolen bases) and a way to track that "small ball" is a winning strategy particularly when it counts - in the post season. This appears to be the case given the tenor of the article by Buster Olney of ESPN. My minor issues with productive outs include:
- The definition is a little strange in that it does include pitchers bunting with one out but not position players. It doesn't seem to make sense to exclude very poor hitting position players (or any position players for that matter) since poor hitting is why a pitcher would sacrifice with one out in the first place.
- I'm not sure why moving a runner into scoring position with 1 out is not included. I've seen many managers employ just such a strategy even though it is self-defeating.
- I assume for the analysis below (and for common sense sake) that a batter hitting into a double play with nobody out would not get 2 productive outs for his efforts whether a runner scored or was moved into scoring position or not.
- Ground outs with runners on as well as sacrifice flies are already counted as RBIs and so the most productive outs included in this new statistic are already being accounted for.
Which Outs Are Productive?
The first big problem with outs as a positive statistic is that they generally decrease the run potential (the average number of runs scored after a particular situation obtains within an inning) and/or the odds of scoring a single run - which makes them a negative statistic. Therefore some of the outs labeled as productive outs are not productive in any sense. For example, assume that a runner is on first with nobody out. The run potential in such a situation is .953 (according to the table provided by Tangotiger including data from 1999-2002). If the batter makes a productive out the run potential moves down to .725.
Fine, you say, but doesn't it increase the odds of scoring that single run? Nope. Before the productive out the odds of scoring at least one run was 43.7% (again using a table from Tangotiger). After the out 40.6%. Not exactly productive. Of course, this sort of analysis doesn't hold for all base-out situations. For example, with a man on second and nobody out the run potential is 1.189 and the probability of scoring 63.2%. If a grounder moves the runner to third the run potential goes down to .983 but the probability of scoring increases slightly to 66.2%. So the obvious question becomes, can we decide which outs are truly productive and only count those? Sure.
In order to calculate which outs are actually productive you simply look at the before and after state for the batter making the productive out. As in our previous example, before the productive out the run potential was .953 and after .725 so the run potential was decreased by .228 by the productive out. If the batter drove in the run via a ground out or sacrifice fly a 1 would be added to the run potential of the final state. The same can be done for the probability of scoring at least one run.
The table below calculates the before and after differences for both run potential and the probability of scoring a single run for the situations covered by the definition of productive outs for non-pitchers (for pitchers I'll grant that in most cases it makes sense for them to bunt but that will have to wait for a later post). Negative numbers are bad, positive numbers are good.
| Runners/Outs | Run Potential
|
Prob of Scoring 1 Run |
| 1/0 | -.228
|
-.031 |
| 2/0
|
-.464 | +.030 |
| 12/0
|
-.106 | +.054 |
| 13/0 (to 2nd) | -.437
|
-.181
|
| 13/0 (scores) | -.179
|
1.000 |
| 23/0
|
-.069 | 1.000 |
| 3/1 | +.134
| 1.000 |
| 13/1 | +.102
| 1.000 |
| 23/1
| -.080
| 1.000 |
The two rows in bold therefore represent situations where a productive out (given an average hitter) is actually detrimental to the team by both decreasing the run potential and the probability of scoring. In two other situations, runner on third and one out and first and third one out, the productive out is truly productive (both values are positive). So to be more accurate Elias and ESPN should drop the first pair of situations from their stat and only include the second pair. But what about the four mixed states?
In the third set of situations that include runner on 2nd and nobody out, 1st and 2nd nobody out, first and third nobody out, and 2nd and 3rd one out, the run potential decreases while the odds of scoring a run either increases or is 100% since the run scored. This leads to the second big problem with productive outs.
Outs are Context Sensitive
The actual utility of "productive outs" is heavily dependant on the context in which they are made. For example, a productive out in the first inning of a 0-0 game is far less valuable than the same productive out when your team is tied in the bottom of the 9th. The difference is even greater when compared with a productive out that occurs in a 10-1 game in the 5th inning. Basically, because an out is never the desired outcome, any outs that are made fall on a continuum of negative outcomes. Of course, one could argue along the same lines with any counting statistic such as hits or doubles. Certainly a double is more valuable with men on base than with nobody on. But the difference here is that a double is purely positive and so always has positive value in terms of run potential and scoring probability. With making outs the situation dictates the utility.
To find out that Alex Sanchez is 14 of 20 in productive out situations is all fine and good and it does tell me that he has been effective in moving runners over. But it tells me nothing as to how valuable those outs actually were relative to other outcomes or more importantly how many times he came through with hits ("productive hits") in such situations (I assume a productive out opportunity is only recorded when the batter makes an out).
To illustrate the context sensitive nature of productive outs one can look at a Win Expectancy chart like that provided by Tangotiger in the productive out situations and perform the same calculation we did before (A Win Expectancy table shows the probability of the home team going on to win a game when faced with a particular situation based on outcomes of actual games. For example, the odds of the home team winning when the game is tied in bottom of the 7th with 1 out and a runner on first is 59%) All situations shown in the table below are for the home team.
| Runners/Outs | Bottom 7th Down 1 | Bottom 9th Tied |
| 1/0 | -.032 | -.012 |
| 2/0 | -.022 | +.023 |
| 12/0 | +.001 | +.023 |
| 13/0 (to 2nd) | -.066 | -.090
|
| 13/0 (scores) | +.015 | 1.000 |
| 23/0 | +.036 | 1.000 |
| 3/1 | +.054 | 1.000
|
| 13/1
|
+.061 | 1.000 |
| 23/1 | +.033 | 1.000
|
As you would expect in these situations the productive outs are more productive (especially when tied in the bottom of the 9th of couse), except for with a runner on 1st and nobody out and runners on 1st and 3rd with an out that moves the trailing runner to second without scoring the runner from third. But that is exactly the point - their value fluctuates with the situation. I have no doubt that moving runners along can be important in helping win baseball games at critical junctures. I just don't think it has any meaning when tracked as a general counting statistic.
Is Making Productive Outs a Skill?
Related to the whole point of tracking productive outs is the notion that there is a skill to them. Certainly bunting, which is a part of productive outs, is a skill but do we really think that ground outs that advance or score runners with the infield playing back can be thought of as a skill? Same goes for fly balls hit just deep enough or to the weak throwing outfielder or with the fast runner on third.
As you peruse the list on ESPN it does appear that hitters who strike out less frequently and are good bunters wit good speed are generally higher on the list (Sanchez, Taguchi, Womak, Clayton, Perez, Pierre, Miles). But my guess is that a big part of the reason for this is the fact that these are hitters who can't get the big hit that maximizes runs and so are utilizing their strengths by bunting and slapping the ball in the infield. Looked at in this light productive outs are a negative statistic. This leads us to the final criticism.
Do Productive Outs Correlate with Winning?
In the final analysis productive outs should be measured on whether or not they help the team win. As I mentioned previously I believe productive outs are strategic and not general purpose and so a team that has alot of them will - all other things being equal - perform worse than a team that has few.
Earlier this season Larry Mahnken over at Hardball Times did an analysis of this question that backed this up and found that:
- "The 15 teams above average in POP have scored 4.77 runs per game, have a .334 POP and a .468 winning percentage.
- The 15 teams below average in POP have scored 4.74 runs per game, have a .270 POP and a .532 winning percentage.
- The top five teams in POP have scored 4.33 runs per game, have a .351 POP and a .392 winning percentage.
- The bottom five teams in POP have scored 4.74 runs per game, have a .230 POP and a .534 winning percentage."
More to the point, he found that the correlation of productive out percentage to wins was only .476. On the contrary, a few years back Stats, Inc. did a study of games from 1993-1997 that showed that teams with the higher OPS (on base plus slugging percentage, a quick and accurate sabermetric measure) in games won those games 85.2% of the time.
But you say, what about productive outs in the post season? There the run environment is contrained and so the value of productive outs should be greater. This was the contention in the original article on ESPN as well. In that article Olney discussed how the Marlins won last year's World Series largely because they led in productive outs 9-5 and that in 62.5 percent of the post season series since 1969 the leader in productive outs has gone on to win the series.
From a purely strategic perspective Kevin Mello and Vijay Mehotra did a study published in the February issue of By The Numbers on run potential and scoring probabilities in all World Series games from 1982-2002. Their conclusion is that one-run strategies are generally not more helpful in the World Series than in the regular season. However, using their table I calculated the following run potential and scoring probabilities for productive outs.
| Runners/Outs | Run Potential
|
Prob of Scoring 1 Run |
| 1/0 | -.018 | +.08 |
| 2/0
|
-.160 | +.053 |
| 12/0
|
-.400 | +.030 |
| 13/0 (to 2nd) | -.050
|
-.147 |
| 13/0 (scores) | -.240
|
1.000 |
| 23/0
|
+.110 | 1.000 |
| 3/1 | +.170
| 1.000 |
| 13/1 | +.130
| 1.000 |
| 23/1
| -.420
| 1.000 |
As you can see, the value of productive outs is higher in the World Series because the run environment is contrained (they found that scoring drops about 13% in the World Series) partially because of the better pitching but also in part because teams tend to play one-run strategies in the post-season. This has the effect of turning the runner on first and nobody out situation into a slightly positive situation for scoring a run and (strangely) the second and third nobody out situation into a positive for maximizing runs. The general rule holds, however, that giving up the first out is a bad play if your aim is to maxmize scoring.
So what of the argument that the 9-5 advantage in productive outs by the Marlins led to their 6 game victory in 2003? If one looks a little deeper at the 2003 World Series you quickly discover that it likely wasn't that the Marlins small-ball strategy paid off in a big way but rather that the Yankees failed to get hits with runners on base.
First, the Marlins scored 17 runs in the series and the Yankees 22. Using Bill James Run Created formula (RC = (H+BB+HB-CS-GIDP)*(TB+.26*(BB-IBB+HB)+.52*(SB+SF+SH))/(AB+BB+HB+SF+SH) you find that the Marlins should have scored 16.9 runs and the Yankees 27.7. The Marlins hit their mark while the Yankees underperformed by almost a run per game. A quick look at retrosheet reveals that the Marlins hit .232 for the series, .214 with runners on base, and .233 with runners in scoring position. The Yankees on the other hand hit .261 in the series, .182 with runners on base, and .140 (7 for 50) with runners in scoring position. Over a span of just six games, three of which were lost 3-2, 4-3, and 6-4, that pretty well accounts for the outcome of the series.
My guess is that the correlation between winning in the post season and the number of productive outs referenced by Olney simply reflects the fact that winning teams get more runners on base, have more productive out opportunities, and therefore usually end up making more productive outs. The 2003 World Series was likely an exception. A slog through the play-by-play files found on retrosheet should definitively answer the question. A project for another day.
Conclusion
So what does it all mean? Productive outs are an interesting idea but one that hasn't been fully thought out by its creators. Rating players by their Productive Out Percentage (POP) is likely meaningless because the values of the outs are so situationally dependant. If it is going to be tracked it should at the least be modified to exclude:
- The runner on first and nobody out situation
- The runners on first and second and nobody out where the out does not score the run
However, I predict that after a brief interest it will fade away much like Number of Bases Touched in the 1870s and Game Winning RBIs from the 1980s because it does not correlate to winning.
Making productive outs is not a general purpose skill but rather a strategy to be employed very narrowly. As a result it does not have the "house advantage" that new Dodger GM Paul DePodesta wrote about in an article I discussed earlier this year:
"I was on a quest to find relevant relationships. Usually it wasn't as simple as 'if X then Y.' I was looking for probabilistic relationships. I christened the new model in the front office: 'be the house.' Every season we play 162 games. Individual players amass over 600 plate appearances. Starting pitchers face 1,000 hitters. We have plenty of sample size. I encouraged everyone to think of the house advantage in everything we did. We may not always be right but we'd be right a lot more often than we'd be wrong. In baseball, if you win about 60% of your games, you're probably in the playoffs."
Posted by
Dan Agonistes
at
5:00 AM
1 comments
Saturday, August 07, 2004
The Four Man Rotation Redux
On my recent vacation in Florida I had a chance to pick up The Neyer/James Guide to Pitchers. Wow! In addition to providing an entertaining (if brief) history of pitching followed by histories of each pitch and biographies of some underrated pitchers Rob Neyer and Bill James provide an encyclopedia of what pitchers threw. So, for example if you want to know what Fergie Jenkins threw you can flip to his entry and see the following:
"Pitches: 1. Four-Seam Fastball 2. Two-Seam Fastball 3. Slider 4. Curve 5. Forkball (as change)"
They then document the source of this info along with quotes from other sources about the style or way in which the pitcher threw. Absolutely fascinating - a must have for those interested in baseball history and particularly for those who love pitching. . BTW, I see that Neyer has now published some errata for the book.
Pitcher Abuse?
The only controversial aspect of the book and topic that has been hot in the sabermetric community as of late is an essay by James called "Abuse and Durability" and a rebuttal by Rany Jazayerli and Keith Woolner from Baseball Prospectus titled "A Response in Defense of PAP" (I blogged briefly about PAP (pitcher abuse points) recently). These essays reflect James' criticism of the methodology used by Jazayerli and Woolner and their response. In short, James uses a set of comparison studies to compare pitchers marked as abused by PAP against other pitchers and tries to determine how the value of each set of pitchers changes over time. He concludes that the "pitchers that Jazayerli and Woolner identify as 'abused' declined in value by 10% in the following season. That can't possibly be an unnatural decline. This system simply does not identify at-risk pitchers, period." [emphasis in the original]
This idea that pitchers could be abused and the modern emphasis on pitch counts can be traced to the work of one-time sabermetrician for the Texas Rangers Craig Wright in his book The Diamond Appraised as well as to James himself in the early Baseball Abstracts (in 1982 he discusses Billy Martin's pattern of abuse). In fact, James quotes Wright's book in criticizing PAP. I pulled out my copy (no sabermetric library should be without one) and recalled that Wright did his study based on what he called Estimated Batters Faced Per Start (BFS) calculated as:
BFS = ((IP * 3) + H + BB) / ((GS + ((G-GS) * .5)))
Wright then identified those pitchers with high BFS numbers at young ages to see how their careers progressed versus other pitchers with low BFS values at similar ages (typically 18-24). After doing his comparison study on pitchers with high BFS numbers in the teens versus those who were not worked as hard Wright lays out the following guidelines for pitcher use (p211):
"1. As a teenager, a pitcher should not be allowed to throw two-hundred innings seasons or have a BFS over 28.5 in any significant span (150-plus innings)."
2. A teenage pitcher should not start on three days rest..."
3. For ages twenty to twenty-two, they should average no more than 105 pitches per start for the season (105 pitches is the rough equivalent of a 30.0 BFS). A single game ceiling should be set at 130 pitches.
4. For age twenty-three to twenty-four, the restraints can be eased up, but their season average should stay under 110 pitches in most cases. The single game ceiling can be jumped to 140 as long as the pitcher is still strong."
To Wright it is clear that the age of the pitchers is the crucial element coupled with the average number of pitches over a prolonged period as well as the maximum number of pitches in a start. It is this crucial fact that James points out when he says "Craig was concerned mostly with the effect of hard usage on young pitchers." [emphasis in the original]. Although James calls the omission of a strong age component a "marginal concern" in the PAP system it appears to me to be the central reason why James doesn't detect abuse. His studies simply look at the list of pitchers selected as most abused by the PAP system which consist of pitchers of all ages, many of whom since they are mature, aren't really impacted by high pitch counts as discussed in this interesting article on mlb.com by Will Carroll.
Adjustments for age have been made to PAP since 2001 as pointed out in the rebuttal but James downplays their importance in his criticism. Woolner and Jazayerli, however, in defending their work say that "Bill is almost certainly right when he states that almost all of the risk inherent in throwing too many pitches occurs in a pitcher's formative years. The connection between overuse and catastrophic injury seems to drop off quickly past age 25 or so." The PAP system, however, still includes all pitchers and so likely "labels" some pitchers as abused who are mature and can handle the work load.
Checking For Abuse
Because Wright performed his study in the late 1980s I wanted to check and see if his recommendations had been followed so I ran some queries to test points 1,3, and 4.
To test point (1) I ran a query showing those teenage pitchers with 200 inning seasons or a BFS over 28.5 in 150 innings or more since 1960 (of course now that batters faced is known I did not need to rely on Wright's formula). The only two pitchers that match these criteria since 1960 were a pair of 19 year olds - Gary Nolan for the Reds in 1967 (226 IP and BFS of 29.1) and Wally Bunker in 1964 for the Orioles (214 IP, 29.3 BFS). Nolan's arm injury is well known although Bunker also developed a sore arm in 1965 and was never the same winning only 41 more big league games. The sample size is too small to say anything definitive of course.
To test point (3) I ran a query showing those pitchers who threw more than 105 pitches per start using Tangotiger's basic pitch count estimator (3.3*((IP * 3) + H + BB))+(1.5*SO)+(2.2*BB) but replacing (IP*3)+H+BB with actual batters faced) at ages 20 through 22 (as of July 1) in over 30 starts since 1960. There were 51 seasons that qualified.
Name Year Age GS IP BFS P/GS
--------------- ------ ----------- ------ ------------ ------- -------
Jose Rijo 1986 21 26 193.7 32.9 127.9
Mark Fidrych 1976 21 29 250.3 34.3 122.4
Bert Blyleven 1973 22 40 325.0 33.0 122.3
Frank Tanana 1975 21 33 257.3 31.2 120.0
Frank Tanana 1974 20 35 268.7 32.2 118.8
Vida Blue 1971 21 39 312.0 30.9 118.7
Dwight Gooden 1986 21 33 250.0 30.9 116.4
Don Gullett 1973 22 30 228.3 31.4 116.3
Dwight Gooden 1985 20 35 276.7 30.4 116.2
Jon Matlack 1972 22 32 244.0 31.3 116.2
Fernando Valenz 1982 21 37 285.0 31.2 116.1
Mike Nagy 1969 21 28 196.7 31.1 115.4
Denny McLain 1965 21 29 220.3 30.4 114.9
Larry Dierker 1968 21 32 233.7 30.7 114.8
Ray Sadecki 1961 20 31 222.7 30.6 113.8
Bert Blyleven 1972 21 38 287.3 30.5 113.6
Greg Maddux 1988 22 34 249.0 30.8 113.0
Catfish Hunter 1967 21 35 259.7 30.1 113.0
Jerry Garvin 1977 21 34 244.7 30.8 112.6
Dean Chance 1963 22 35 248.0 30.2 112.6
Dave Rozema 1977 20 28 218.3 31.8 112.5
Dave Boswell 1967 22 32 222.7 28.8 112.1
Floyd Youmans 1986 22 32 219.0 28.3 110.9
Ismael Valdes 1995 21 27 197.7 29.8 110.8
Britt Burns 1980 21 32 238.0 30.3 110.6
John Smoltz 1989 22 29 208.0 29.2 110.5
Bert Blyleven 1971 20 38 278.3 29.6 110.0
Bill Butler 1969 22 29 193.7 28.8 109.9
Burt Hooton 1972 22 31 218.3 29.5 109.6
Ramon Martinez 1990 22 33 234.3 28.8 109.6
Roger Erickson 1978 21 37 265.7 30.3 109.5
Jim Palmer 1966 20 30 208.3 28.9 109.4
Kerry Wood 1998 21 26 166.7 26.9 109.4
Joe Coleman 1969 22 36 247.7 28.9 109.1
Ed Correa 1986 20 32 202.3 27.7 108.9
Ray Culp 1963 21 30 203.3 28.0 108.7
Mark Lemongello 1977 21 30 214.7 30.3 108.1
Carlos Zambrano 2003 22 32 214.0 28.3 107.9
Denny McLain 1966 22 38 264.3 28.4 107.4
Steve Dunning 1971 22 29 184.0 28.0 107.4
Gary Nolan 1970 22 37 250.7 28.4 106.9
Don Robinson 1978 21 32 228.3 29.3 106.9
Clay Kirby 1970 22 34 214.7 27.9 106.7
Wayne Simpson 1970 21 26 176.0 28.1 106.4
Dave Rozema 1978 21 28 209.3 30.3 106.2
Bret Saberhagen 1985 21 32 235.3 29.1 106.0
Catfish Hunter 1968 22 34 234.0 28.5 106.0
Dennis Eckersle 1976 21 30 199.3 27.4 106.0
Storm Davis 1983 21 29 200.3 28.7 105.9
Les Cain 1970 22 29 180.7 27.3 105.7
In these 50 seasons are 42 unique pitchers. It is telling that only four pitchers have made the list since 1990 which is possibly an indication that organizations are more careful and indeed are heeding the advice of Wright at least indirectly. Of course, the things that jumps out to me are the two recent additions to this list, Kerry Wood and Carlos Zambrano of the Cubs. Wood did incur a serious injury that wiped out the end of 1998 and all of 1999 and suffered a sore tricep this season. Zambrano has yet to be injured but his precense on the list and his continually high pitch counts make me a little uneasy. (As an aside it should be noted that the pitch count estimator estimated 107.9 pitches per start for Zambrano in 2003 when he actually had 106.4).
To test point (4) I ran a query of those pitchers aged 23 or 24 who threw more than 115 pitches per start in 25 starts or more. This produced a list of 36 seasons with 30 unique pitchers.
Name Year Age GS IP BFS P/GS
--------------- ------ ----------- ------ ------------ ------- -------
Rick Sutcliffe 1979 23 30 242.0 33.9 124.7
Jim Bouton 1963 24 30 249.3 33.5 124.2
Fernando Valenz 1984 23 34 261.0 31.7 122.1
Jim Hughes 1975 23 34 249.7 32.4 120.7
Dick Drago 1969 24 26 200.7 32.6 119.4
Bert Blyleven 1975 24 35 275.7 31.5 119.4
Joe Coleman 1970 23 29 218.7 31.7 119.2
Jack Hamilton 1962 23 26 182.0 31.5 119.0
Denny Lemaster 1963 24 31 237.0 31.5 119.0
Jim Maloney 1963 23 33 250.3 30.6 118.9
Tommy Greene 1991 24 27 207.7 31.7 118.7
Lynn McGlothen 1974 24 31 237.3 31.9 118.4
Jim Merritt 1967 23 28 227.7 32.5 118.4
Livan Hernandez 1998 23 33 234.3 31.5 118.3
Jon Matlack 1974 24 34 265.3 31.6 118.0
Frank Tanana 1977 23 31 241.3 31.4 117.8
Fernando Valenz 1983 22 35 257.0 31.3 117.5
Mel Stottlemyre 1965 23 37 291.0 32.1 117.5
Clay Kirby 1971 23 36 267.3 30.8 117.4
Don Sutton 1968 23 27 207.7 31.4 117.3
Denny McLain 1968 24 41 336.0 31.4 117.3
Dean Chance 1964 23 35 278.3 31.2 117.3
Bert Blyleven 1974 23 37 281.0 31.1 117.2
Dennis Eckersle 1978 23 35 268.3 32.0 117.1
Andy Messersmit 1969 23 33 250.0 30.5 117.1
Jim Kaat 1962 23 35 269.0 31.8 117.0
Joe Coleman 1971 24 38 286.0 30.9 116.8
Larry Dierker 1970 23 36 269.7 31.4 116.7
Dave Righetti 1982 23 27 183.0 29.8 116.1
Steve Barber 1961 23 34 248.3 30.6 116.0
Bret Saberhagen 1987 23 33 257.0 31.8 115.7
Chuck Estrada 1961 23 31 212.0 29.8 115.5
Dick Ellsworth 1963 23 37 290.7 31.4 115.4
Stan Bahnsen 1968 23 34 267.3 31.5 115.4
Sam McDowell 1966 23 28 194.3 28.8 115.1
Blue Moon Odom 1968 23 31 231.3 30.6 115.0
Once again very recent pitchers don't really show up on the list with the exception of Livan Hernandez in 1998.
In looking at these lists it is apparent that something has changed in the workloads of young pitchers since the late 60s and 70s. I believe there are three primary reasons:
These last two points are illustrated by the fact that if you look at the top 100 seasons since 1960 ranking them in descending order by IP per start you won't find any seasons after 1989. In fact, it isn't until 147th place that Roger Clemens' 1997 season appears when he averaged 7.8 innings in 34 starts.
But the real question is whether this altered pattern has actually decreased the number of injuries to starting pitchers. Jazayerli and Woolner suggest that the jury is still out but are hopeful in their rebuttal stating that "A revolution in the management of starting pitchers is underway, and the early signs suggest that the revolution may well lead to fewer injuries." They don't, however, list what these signs are. This question is also difficult to analyze since teams use the disabled list more frequently today than in the past in part because diagnostic technology has improved and therefore precautionary DL stints are more common, and partly because of the large investment that players represent that encourages a better-safe-than-sorry approach.
I think all parties agree that its clear that abuse along the lines documented by Wright indicates that young pitchers can and were overworked fairly routinely in the past which often led to serious injuries. However, if evidence of reduced injuries to young starters in the present era fails to materialize, it could well be that James is correct in surmising that most injuries to pitchers both young and old are catastrophic and not the result of overuse. The truth probably lies in the middle somewhere - young pitchers are prone to both overuse and catastrophic injuries while mature pitchers typically incur catastrophic injuries. Understanding the relative frequency of each is the key to formulating studies and systems that might detect it.
A second explanation as to why recent pitchers don't show up on these lists assuming that injury frequency has not lessened, might be the alteration in workload from a very young age. Since young pitchers have been so closely monitored throughout their careers from little league on up, their arms are not as strong as pitchers of the past. In that case the prevelance of arm injuries might not change but would simply be triggered at lower and lower pitch counts. For example, in the 1970s a game of 140 pitches might be considered the borderline for a mature starter whereas today it might be 120 pitches and tomorrow 100 pitches.
Four Man Rotation
And finally, this leads us to the title of this post. Could a return to the four man rotation actually help the situation? Wright found in his study that pitchers are not less effective on three days rest than on four when looking at data from 1986-87 and advocated a return to the four-man rotation. This was backed up by a study done by Jazayerli using data from 1978 to 1995. Of course, far fewer pitchers ever pitch on three days rest today. The last manager to be a strong proponent was Earl Weaver in the early 1980s and the last experiment with it I can remember is the Reds Bob Boone's in 2003 which ended by mid season despite support from pitching coach Don Gullet and assistant GM Brad Kullman.
It seems to me (and others I've read) that the focus on reducing pitch counts has become married to the assumption that more rest between starts is also good for young pitchers. You could easily see where this might be the case since both reduce the number of pitches a pitcher throws over the course of a seaon. But is it just possible that for most pitchers throwing fewer pitches (meaning a maximum of 105 for young pitchers and 120 or so for mature pitchers for example) more frequently (every four days instead of five) will actually build arm strength, endurance, and lead to fewer injuries in the long run? Of course it will also have the side benefit of allowing better pitchers to throw more innings thereby producing more wins for the team that employs the four-man rotation. I'd love to see some team try it in 2005 but now that the pendulum has swung in the opposite direction it will likely take a while for it to swing back.
Other reading on pitch counts:
Kicking the crutches out
What Pitch Counts Hath Wrought
What Pitch Counts Hath Wrought: Part Deux
Swinging from the Heels: Painting a Fake Tunnel on a Blind Alley
Posted by
Dan Agonistes
at
5:00 AM
0
comments
Rejoice and Be Glad
Statistics are the lifeblood of baseball. The true fan cannot discuss baseball without referring to them passionately and often. For many Americans the morning ritual of reading the box scores is a momentary escape and an assurance that all is well with the world. The mere mention of numbers like 73, .406, 755 and 56 are powerful enough to evoke images of ballplayers past and present. The emphasis on statistics throughout baseball history has been one of the truly fortunate (or perhaps divinely inspired?) aspects of the game. The ability to now, with data spanning back well over a century, look back over the stat lines of players long since dead and get a feel for the kind of player they were connects us to that past and helps give baseball its sense of continuity that makes it the National Pastime. Truly, as noted sabermetrician (the word that describes those who work to analyze the stats, the word derived from the acronym of the Society for American Baseball Research or SABR, www.sabr.org, to which I proudly belong) Bill James has said, we love baseball statistics because they have acquired the power of language.
Statistics of course are also what fuels the fire of the never ending controversies of who was better than who (or rather whom). Without the baseline of career batting average, how could one even argue whether or not Shoeless Joe Jackson - a career .356 hitter - was better than Ted Williams at .344? Or whether or not Hank Aaron's 755 homeruns were a more impressive feat than Babe Ruth's 714 or Barry Bonds’ 685 and counting? These players never played in the same game and were separated by decades. And although the statistics haven’t always been optimally fashioned or tracked (a situation that the likes of Retrosheet, www.retrosheet.org are attempting to remedy), they do give us a place to begin the discussion.
To understand the depth and richness that statistics provide to the game, one need only reflect on the careers of two players, Josh Gibson and Satchell Paige. Both in their day were considered the best players in the Negro Leagues. Because statistics were rarely compiled and the level of competition varied, it is now impossible to pin down with any certainty how great these players were. Dizzy Dean, one of the best pitchers in baseball in the 1930's considered Paige one of the fastest and best he ever saw. A faint image of Paige’s greatness can be detected in the fact that he pitched effectively in the major leagues from 1948 to 1953 in his mid to late forties (Satchell never seemed to be able to make up his mind whether he was born in 1906 or some other year in the vicinity). Had he played in the majors in his prime he may have racked up enough wins and strikeouts to rank with the greatest in the game. Gibson's power was legendary and he was said to have hit over 80 homeruns in a Negro League season. Could he have challenged or even beaten Ruth? These questions cannot be answered with any degree of certainty or precision at all since there is no baseline to begin the discussion.
In another sense baseball statistics are meaningful because baseball is a team game of individual accomplishments and confrontations. The individual's performance is paramount at any one moment and can be separated from the team's performance. No doubt, what the individual does has a great impact on the success or failure of his team, but what he does is also in a sense separate from the team as well. Statistics also lend themselves to baseball because baseball, in the jargon of my profession as a computer scientist, is "event-driven". By "event-driven" I mean that a baseball game unfolds as a series of discrete events which can be separated and analyzed. A pitch is thrown, what are results? Was it a strike? Did the batter swing? What was the count? Did the runner move? Who field the ball? That event can then be recorded for posterity and tracked (I might add with increasing precision here at MLB.com).
In other games, notably basketball and football, individuals act in concert with teammates to attempt to produce desirable results. In these games it is not so easy to separate the actions of one man from his team. In a typical play in football, 22 players are moving simultaneously and the outcome of the play is dependent on a whole host of variables that are not easily identified. When the play is over what can you record? The ball moved from the 20 yard line to the 23 yard line and Priest Holmes carried the ball before being tackled by Brian Urlacher. While that accounts for two of the 22 players on the field, what do you say about the other 20? Certainly they were an integral part of how the play unfolded and yet there is no easy way to evaluate their performance. In contrast, the confrontation of batter versus pitcher and even fielder versus ball (despite Branch Rickey’s protestation that “There is nothing on earth anybody can do with fielding”) are much more easily recorded and analyzed.
And finally of course, statistics are the glue that binds generations of baseball fans and make baseball unique in its role as a heritage passed down from fathers to sons (or daughters in my case). When a father first tells his son about DiMaggio and 56 or Williams and .406 or when a father and daughter witness George Brett’s 3,000th hit they are speaking a language that connects them to each other and to an ongoing story.
So today as you peruse the box scores take a second to reflect on the fortuitousness before you and along with the Psalmist “rejoice and be glad in it.”
(this is a revision of an essay I previoulsy posted on this blog)
Posted by
Dan Agonistes
at
2:08 AM
1 comments
Wednesday, August 04, 2004
Where have the .400 hitters gone?
This is a question that comes up from time to time among baseball fans who lament that the giants of yesteryear - Ty Cobb, Rogers Hornsby, and Ted Williams among them are long gone and that we'll never see their kind again.
If one doesn't accept the premise that these immortals were truly superior to modern stars such as George Brett, Tony Gwynn, and Ichiro Suzuki what accounts for the disappearance of the .400 hitter? Some of the factors that have been discussed in recent years that might cause modern hitters to be at a disadvantage include changes in the game itself - the increase of pitching specialists, the development of the slider and split finger fastball, the prevalence of night games, and intracontinental travel. But do any of these hold water?
Recently, this topic came up again on the SABR listserv in the context of sabermetric references made by the late Harvard paleontologist and baseball fan Stephen Jay Gould in his book Triumph and Tragedy in Mudville: A Lifelong Passion for Baseball. The discussion revolved around an essay that Gould wrote for Discover magazine in 1986 and that was reprinted both in his 1996 book Full House and Triumph and Tragedy under the title "Why No One Hits .400 Any More?"
In epitome Gould's argument was that .400 hitters haven't disappeared because of cosmetic changes in the game or that the heroes of the past were supermen, but rather as the natural consequence of of an increasing level of play that comes closer to the "right-wall" of human ability coupled with stabilization of the game itself. These factors tend to decrease the differences between average and stellar performers. As a result, although the mean batting average has remained roughly .260 since the 1940s, there are now fewer players at both the left and right ends of the spectrum. In other words, "variation in batting averages must decrease as improving play eliminates the rough edges that great players could exploit, and as average performance moves towards the limits of human possibility and compresses great players into an ever decreasing space between average play and the immovable right wall."
To support his argument Gould (actually his research assistant) calculated the standard deviation (a measure of the spread in values assuming a normal distribution) of batting averages over time and plotted them in a graph and presented the following table showing also the coefficient of variation (the standard deviation divided by the mean useful for comparing distributions with different means).
Decade Stdev Coeff
1870s .0496 19.25
1880s .0460 18.45
1890s .0436 15.60
1900s .0386 14.97
1910s .0371 13.97
1920s .0374 12.70
1930s .0340 12.00
1940s .0326 12.23
1950s .0325 12.25
1960s .0316 12.31
1970s .0317 12.13
This certainly shows a trend towards decreasing variability over time. Gould's conclusion was that this decreasing variability was due to refinments in the game (standardized techniques for pitching starting in the 1880s, the introduction of gloves, stabilization of the number of balls and strikes, and refinement of strategies) coupled with the entire system moving farther towards the limits of human ability in much the same way that track athletes move closer to that wall with each Olympics thereby decreasing the variation in sprint times. However, since Gould's data was published almost 20 years ago I decided to take a look and see if I could reproduce Gould's data and add data for the last two plus decades.
To do so I used the Lahman database and calculated the league batting average and OPS for the 250 league seasons included in the database. I then selected the 74,277 seasons where a player batted at least once calculating their batting average, OPS, and plate appearances along with their respective league averages. Finally, I selected all of the players with more than 2 at bats per game (relative to the league schedule) which pared the list down to 18,104 seasons. From these I calculated the standard deviation (using the league average) and coefficient of variation by decade and produced the following table.
Decade Seasons Stdev Coeff
1870s 587 .0508 18.60
1880s 1189 .0423 16.85
1890s 1004 .0402 14.57
1900s 1110 .0373 14.68
1910s 1220 .0372 14.56
1920s 1194 .0369 12.93
1930s 1233 .0349 12.53
1940s 1149 .0329 12.64
1950s 1145 .0334 12.88
1960s 1448 .0319 12.83
1970s 1867 .0316 12.33
1980s 1970 .0299 11.54
1990s 2103 .0311 11.73
2000s 885 .0310 11.71
As you can see I wasn't able to recreate Gould's results precisely for some reason but got very close in several decades including the 1970s (.317 to .316) and the 1910s (.371 to .372). Overall there appears to be less spread in my data in the early years and more in the latter years, again for unknown reasons. I've tried several different cutoffs (200 at bats, 250 at bats, etc.) but none have come any closer to reproducing Gould's numbers. It is also possible that Gould used a different dataset that was less complete for the years prior to 1900. The Lahman database includes the National Association for 1871-1875 and the American Association for 1882-1891, the Player's League for 1890 and the Federal League for 1914-1915. Adding more player seasons will tend to decrease the standard deviation. Overall though, the trend towards lower standard deviations seems to continue with the addition of the 1980s through 2003 as the three lowest standard deviations and coefficients of variation are for those three decades.
I then produced the following scatter plot that shows the standard deviations for each year.
Intuitively, it seems as if Gould's argument holds. However, several suggestions and points of discussion that were brought out on the SABR-L list included:
- The analysis shows that standard deviations have fallen over time but much less so since the 1940s. Many then agreed that the stabilization of the game had occurred by the 1940s.
- In order to test Gould's hypothesis some argued that higher standard deviations should be found for expansion years since the talent pool expands letting in more players, some of whom would not have been in the major leagues the previous year. When looking at the years 1901, 1961, 1962, 1969, 1977, and 1993 there is no evidence that the standard deviations were greater in these years. The reason this study doesn't find that result may be because when looking at players with 2 at bats a game or more in an expansion year you're really looking at players who were already bonafide major leaguers but who simply didn't get the at bats before expansion. When lowering the cutoff to 50 at bats you do see small increases the stdev in 1901, 1962, 1969, 1977, and 1993.
- The league average has not been consistently .260 and so some argued that in order to perform the calculation you need to standardize the averages. I reran the numbers computing the average as AVG/lgAVG*0.260 and produced the following table. These results are very similar to the first set although the variation appears more consistent from the 1920s through the 1970s before dropping in the 1980s and later:
Decade # Stdev
1870s 587 .0483
1880s 1189 .0439
1890s 1004 .0379
1900s 1110 .0381
1910s 1220 .0379
1920s 1194 .0336
1930s 1233 .0325
1940s 1149 .0328
1950s 1145 .0335
1960s 1448 .0334
1970s 1867 .0321
1980s 1970 .0300
1990s 2103 .0305
2000s 885 .0305
- To be more precise some argued that you should weight the averages by the number of at bats. However, since we're already selecting those who garnered significant playing time with greater than 2 at bats per game (and I'm too lazy to do the weighted calculation) I doubt weighting would change the results very much.
- Perhaps the disappearance of the .400 hitter has more to do with an emphasis on power over average since the 1920s as some speculated. In other words, hitters are knowingly sacrificing average for power in the modern era. I have no doubt that generally this is true as both the increase in strikeouts and the increase in the diversity of skills of batting champions shows. I'm just not sure that there hasn't always been a substantial population of players who have focused on average and who would test the limits of singles hitting. In addition, major league baseball and the general public still hold batting average in high regard and so players are still rewarded for high averages over on base percentage.
- Others contended that the stdev of other measures such as OPS and SLUG have increased over time or held steady and so would put a hole in Gould's hypothesis. I don't think that's the case since clearly the increase in slugging percentage (for example after 1920) doesn't reflect on the talent level but rather on many players adopting a different style of play coupled with rule change which would naturally increase the standard deviation. On a side note some mentioned that Gould's analysis was fatally flawed since it takes into consideration only batting average, a dubious although ubiquitous, measure of offensive value. I wouldn't disagree if the discussion was about pure offensive value. The question however, is what happened to .400 hitters, which by definition looks at batting average.
- Other argued that in order to test Gould's hypothesis you should really be looking at the percentage of players several standard deviations above the league average since the players in the population are not a normal distribution but rather the right hand tail of the distribution (since players much below the league average won't get enough at bats while players above the league average will). I performed this calculation selecting only those players who had a batting average greater than league average plus 2.5 times the standard deviation. The percentage of players per decade is shown below. Interestingly, one would expect a higher percentage of players in the early years although this is not the case. The percentage begins to drop only in the 1960s. Why this is the case I don't know.
Decade %Players #Players
1870s 0.0153 26
1880s 0.0162 86
1890s 0.0127 112
1900s 0.0172 153
1910s 0.0181 248
1920s 0.0167 226
1930s 0.0161 166
1940s 0.0174 261
1950s 0.0177 236
1960s 0.0138 322
1970s 0.0110 311
1980s 0.0108 328
1990s 0.0097 407
2000s 0.0091 198
- Finally, another way to look at this problem as pointed out by Bill James is to calculate the difference in standard deviations from .400 for the various decades. To do this I took the players with greater than 2 at bats per game and calculated their average batting average by decade. Then I subtracted the average from .400 and divided by the standard deviation for the decade to figure out how many standard deviations the average player in the study was away from .400 during that time period. This analysis showed that indeed players in the 1920s who qualified hit .300 with a league standard deviation of .0369 which puts them only 2.7 standard deviations away from .400. On the contrary in the 1990s the players that qualified hit .276 with a stdev of .0311 putting them 3.98 standard deviations away from .400. The higher average combined with higher standard deviations made it statistically more likely that someone would hit .400 as evidenced by the number of .400 hitters per decade (7 in the 1920s, 0 in the 1990s). It appears from this that both factors, the higher relative averages and the increased standard deviations, were in play to account for the prevalence of .400 hitters. What this also shows is that since league averages have risen in the past 11 years the odds are now slightly better that someone will hit .400 than they were in the 1970s and 80s (although not as high as in the 1930s through 1950s).
Decade # .400 hitters AVG Stdev Stdev from .400
1870s 587 9 0.275 0.0505 2.466
1880s 1189 3 0.261 0.0423 3.283
1890s 1004 11 0.286 0.0402 2.837
1900s 1110 1 0.268 0.0373 3.544
1910s 1220 3 0.272 0.0372 3.445
1920s 1194 7 0.300 0.0369 2.705
1930s 1233 1 0.294 0.0349 3.050
1940s 1149 1 0.275 0.0329 3.800
1950s 1145 0 0.275 0.0334 3.733
1960s 1448 0 0.265 0.0319 4.218
1970s 1867 0 0.269 0.0316 4.138
1980s 1970 0 0.269 0.0299 4.369
1990s 2103 0 0.276 0.0311 3.983
2000s 885 0 0.278 0.0310 3.941
Conclusion
So in the final analysis what does it all mean? My tentative conclusions are:
- Based on intuition and observation Gould was certainly correct that baseball players are better than they were in the past and that the game is better played now than ever before. Analogs with basketball and track make this apparent.
- Gould was partially correct that decreasing standard deviations do in fact record an increasing and standardized level of play. However, decreased league batting averages also played a role in the disappearance of the .400 hitter.
For other analyses on this question see this article and this one.
Posted by
Dan Agonistes
at
6:00 AM
127
comments
Monday, August 02, 2004
Moves at the Trade Deadline
Both the Cubs and Royals were involved in the last minutes moves at the trade deadline last weekend. In short:
- The Royals acquired 22 year-old catcher Justin Huber from the Mets organization in exchange for Jose Bautista whom they acquired earlier this season from Baltimore. Huber will go to AAA and the Royals like his power potential (15 homeruns in A and AA last year). He's also seems to have a pretty good eye putting up a .360 OBP last season while hitting around .270. The Royals want to work on his defense (particularly his throwing) but he could end up being a better bat than John Buck, who has looked overmatched at the major league level thus far. He may also be moved to a corner outfield spot or first base (as if the Royals don't have enough of those guys already).
- The Royals also acquired 27 year-old outfielder Abraham Nunez from Florida (not the one from Pittsburgh) and sent reliever Rudy Seanez packing. Nunez was one who had 2 years added to his age recently and so has lost some of his luster as a prospect. He's also had injuries (hamstring in 2003) that have limited him. He hit .311/.398/.547 for the Albuquerque Isotopes (one of the worst team names you can find by the way) last season but in a notorious hitter friendly environment. He has some pop and has shown a little patience and speed in the past as well. Not sure what the Royals plans are for him with the combination of Mateo, Brown, and Stairs in the outfield. At this point I'd like to see the Royals bench Stairs and see if the other three can play. DeJesus seems to be coming around a bit has hit .292/.378/.403 since the break. Strangely, he doesn't seem able to steal bases at all (2 for 5 thus far).
I like both deals since Huber has more upside than Bautista and Nunez than the aging Seanez and it may just turn out that Huber at least will be a regular in the majors some day.
- The Cubs of course acquired Nomar Garciapara from the Red Sox in a multi-team trade giving up Justin Jones a 6'4" 19 year-old left handed pitcher, Brendan Harris, the organization's top infield prospect, Francis Beltran, the 23 year-old top pitching prospect (40 K's in 35 innings with the Cubs this season), and Alex Gonzales (605 OPS). They also acquired outfielder Matt Murton from Boston who was in A ball hitting .301.
The Bottom Line: The Cubs gave up a lot. Three of their top prospects for possibly 2 months of time from Nomar. Given that the Cubs will not win their division this year and are trailing in the wild card race I'm not sure this is such a good deal. Beltran would have looked nice in the rotation next year if Maddux retires or injuries once again plague Wood and Prior with Harris taking the utility or second base job and Jones still a couple years away. My preference would have been to put Grudz at shortstop and leave Walker at second.
Posted by
Dan Agonistes
at
4:13 PM
0
comments
Baseball Book Review
The BusinessOfBaseball web site headed up by Maury Brown of SABR is now open for business. Some great resources for those interested in the dollars and cents of the national pastime are available on the site including the Blue Ribbon Panel report from 2000 and the 2002-2006 Basic Agreement. Yours truly has also written a short review of The Last Commissioner for the site posted in the reading materials section.
Posted by
Dan Agonistes
at
3:24 PM
0
comments
Wednesday, July 28, 2004
Testing for the Network
A whitepaper Jon Box and I wrote for MSDN on testing for the presence of a network connection in the .NET Compact Framework was just published here. Enjoy.
Posted by
Dan Agonistes
at
9:56 PM
1 comments
Friday, July 23, 2004
Triumph and Tragedy in Mudville
This is a collection of 30 plus essays on baseball written by the late Harvard paleontologist Stephen Jay Gould over the last 20 to 30 years. Among them are two original pieces on forms of streetball in New York City when Gould was growing up and a piece discussing the lure of baseball for intellectuals like himself (which he views as purely contingent on time and place).
There is also a long tribute to Gould (he died in May of 2002) in the forward written by David Halberstam (author of Summer of '49). Many of the remaining essays appeared in various places such as the New York Times Review of Books, some of which I've read before but many of which I hadn't seen. They range from book reviews to short eulogies (of Mickey Mantle for example) to essays. One of his most famous is his essay on the disappearance of the .400 hitter originally written in 1986 for Discover magazine. He, like George Will in perhaps my favorite baseball book Men at Work, views it simply as the natural consequence of of an increasing level of play that comes closer to the "right-wall" of human ability coupled with the increasing maturity of the game. This increasing level of play tends to decrease the differences between average and stellar performers. As a result, since the mean batting average has remained roughly .260 since the 1940s, there are fewer players at both the left and right ends of the spectrum. This also tracks very well with a book that came out a couple years ago that rated the greatest hitters (for batting average) of all time through a series of statistics and determined that Tony Gwynn was the greatest.
Gould's writing is always interesting and even though he was a life long Yankees fan, he rightly despised the DH and aluminum bats. Baseball fans will find plenty to like here.
Posted by
Dan Agonistes
at
9:01 PM
0
comments
Tuesday, July 20, 2004
Baseball Encyclopedia Errata
Just found out that errors and omissions from the Baseball Encyclopedia that I bought a few months back can be found here. Some fairly significant things like the omission of serveral player records.
Posted by
Dan Agonistes
at
8:21 AM
0
comments
Sunday, July 18, 2004
Relativity and OPS
My contention in a previous post was that OPS (on base + slug) is a useful measure of offensive production because of its simplicity, comparative ability, and correlative value. However, when ranking the greatest single seasons in OPS only three players, Barry Bonds, Babe Ruth, and Ted Williams, made the top 10 seasons. Could this be the result of some bias in favor of these particular players? For example, one obvious thought that comes to mind is that since homeruns have been flying out of the park at an increased rate since 1993 (“chicks dig the long ball”) it is easier for a player like Barry Bonds playing in the context of an expanded run environment to amass large OPS numbers by increasing their slugging percentages (the average number of plate appearances per homerun in the period 1960-1992 was about 47, from 1993-2003 it was 35, in other words a player with 600 plate appearances would hit about 12 and half homeruns in the 1960-1992 period and almost 17 in the 1993-2003 period). Another thought is that perhaps Ted Williams was inordinately helped by Fenway Park with its Green Monster and short right field line and Babe Ruth by playing in a park that was after all, the house that he built.
Correcting for League and Year
To see if there is some hidden bias here we can first correct for the context by calculating the league average OPS for the 10 seasons in question and then normalizing the individual’s OPS against the league average, a concept first introduced in The Hidden Game of Baseball. For example, in 2001 the National League OPS was 756, a very high number historically. By taking Barry Bonds’ OPS of 1379 and dividing it by the league average (1379/756) we can calculate a Normalized OPS (NOPS) of 1.82, or simply 182 for short. By performing the same calculation with Babe Ruth’s 1920 season (the same raw OPS of 1379 in a league where the average was 730) his NOPS comes out to 189, a little ahead of Bonds. Here are the before and after rankings.
Raw OPS
2002 NL Barry Bonds SFN 1381
2001 NL Barry Bonds SFN 1379
1920 AL Babe Ruth NYA 1379
1921 AL Babe Ruth NYA 1359
1923 AL Babe Ruth NYA 1309
1941 AL Ted Williams BOS 1287
2003 NL Barry Bonds SFN 1278
1927 AL Babe Ruth NYA 1258
1957 AL Ted Williams BOS 1257
1926 AL Babe Ruth NYA 1253
Normalized OPS Raw LgOPS NOPS
1920 AL Babe Ruth NYA 1379 730 189
2002 NL Barry Bonds SFN 1381 741 186
2001 NL Barry Bonds SFN 1379 756 182
1921 AL Babe Ruth NYA 1359 761 179
1957 AL Ted Williams BOS 1257 707 178
1923 AL Babe Ruth NYA 1309 734 178
1941 AL Ted Williams BOS 1287 728 177
2003 NL Barry Bonds SFN 1278 749 171
1926 AL Babe Ruth NYA 1253 739 170
1927 AL Babe Ruth NYA 1258 747 168
Since these ten seasons were first picked because of their raw OPS numbers it’s now appropriate to open up the field and recalculate the single season leaders taking into account their normalized OPS.
Normalized OPS Leaders Raw LgOPS NOPS
1920 AL Babe Ruth NYA 1379 730 189
2002 NL Barry Bonds SFN 1381 741 186
2001 NL Barry Bonds SFN 1379 756 182
1921 AL Babe Ruth NYA 1359 761 179
1923 AL Babe Ruth NYA 1309 734 178
1957 AL Ted Williams BOS 1257 707 178
1941 AL Ted Williams BOS 1287 728 177
2003 NL Barry Bonds SFN 1278 749 171
1926 AL Babe Ruth NYA 1253 739 170
1946 AL Ted Williams BOS 1164 690 169
As you can see the same three hitters still dominate the list, however, the distribution has changed somewhat with Ted Williams garnering another spot for his 1946 season in a league with a low OPS of 690 and Babe Ruth losing his 1927 season when the league put up a fairly high OPS of 747. Williams’ 1957 season now also looks better in this light moving from 9th to 5th place. And most obviously Babe Ruth’s 1920 season now tops the list with an NOPS of 189. So in answer to part of our question we can fairly confidently say that these three hitter’s accomplishments were not inordinately helped by playing in leagues that were hitter’s paradises. In fact, the first player not of this ruling triumvirate to make the list is Mickey Mantle with his 1957 season (NOPS of 167). The only other contemporary player to make the top 20 is Mark McGwire with his famous 1998 season and an NOPS of 166 tied for 15th.
However, hitters that played in extremely low scoring run environments should be greatly helped by normalizing OPS. For example, consider Willie McCovey’s 1969 season and Carl Yastrzemski’s 1967 seasons.
Raw OPS LgOPS NOPS
1969 NL Willie McCovey SFN 1108 686 162
1967 AL Carl Yastrzemski BOS 1040 651 160
Before normalization McCovey in 1969 ranked as tied for 70th all-time with a raw OPS of 1108. After correcting for a league in which pitchers dominated with an OPS of 686 he jumps to tied for 23rd with an NOPS of 162. Even more dramatically Yaz in 1967 moves from tied for 166th to tied for 27th place.
Correcting for Ballpark
But adjusting for the run environment of the league in which a player plays is only part of the context. The park in which the player plays his home games is another significant aspect. Intuitively this makes sense. It seems obvious that Larry Walker gets a boost from playing in Coors Field while Willie McCovey was hurt by playing in Candlestick Park. To take this into account sabermetricians have devoted themselves to calculating “park factors” or “park effects” for each of the major league parks. Historically this has been done by calculating a Batter Park Effect or BPF and a Pitcher Park Effect or PPF for each team. The calculation of these effects as documented The Hidden Game of Baseball involves not only comparing the scoring in each park with the scoring at other parks but also taking into account that there is a "home cooking" bias where batters naturally hit and pitch better at home (a fact well documented in Curve Ball). In addition, the calculation allows for the fact that a team's hitters do not have to face its pitchers and vice versa. The BPF and PPF are expressed as a percentage of the league average, in other words a BPF of 1.06 would mean that the batter's home park gives him a 6% advantage over the league and a PPF of .95 means that the park helps pitchers to the tune of 5%. Although factors can and are calculated for different offensive events, homeruns, doubles, triples, etc. the overall BPF is calculated based on runs scored. Here are the BPF and PPF as calculated for 2003 sorted by BPF.
Team BPF PPF
MON NL 118 116
KCA AL 113 112
COL NL 112 111
ARI NL 111 109
TEX AL 110 109
TOR AL 105 104
BOS AL 105 104
HOU NL 104 103
MIL NL 102 102
MIN AL 102 102
TBA AL 100 100
CIN NL 100 100
NYN NL 99 99
PIT NL 99 99
SFN NL 99 100
CHN NL 99 99
CHA AL 99 99
ATL NL 97 97
SEA AL 97 98
NYA AL 96 97
SLN NL 96 97
PHI NL 95 96
BAL AL 95 96
DET AL 95 95
FLO NL 94 94
LAN NL 93 94
ANA AL 93 94
CLE AL 93 94
OAK AL 93 94
SDN NL 91 92
Fortunately, the BPF and PPF have been calculated and are present in the Lahman database and so in order take into account the home park we simply need to multiply the NOPS by the BPF divided by 1,000. Here are the single season NOPS leaders shown previously re-sorted with a new column for normalized for park effects.
Raw OPS NOPS NOPS/PF
2002 NL Barry Bonds SFN 1381 186 204
2001 NL Barry Bonds SFN 1379 182 200
1920 AL Babe Ruth NYA 1379 189 182
1921 AL Babe Ruth NYA 1359 179 175
1923 AL Babe Ruth NYA 1309 178 175
1941 AL Ted Williams BOS 1287 177 174
2003 NL Barry Bonds SFN 1278 171 173
1926 AL Babe Ruth NYA 1253 170 173
1957 AL Ted Williams BOS 1257 178 168
1946 AL Ted Williams BOS 1164 169 159
So given that Bonds has played in a relatively poor park for hitters (BPFs of 91 in 2001 and 2002) gets helped while Williams is hurt by the high BPFs of Fenway Park that are consistently over 100. And so it is once again appropriate to recreate the top 10 list with park effects.
Raw OPS NOPS NOPS/PF
2002 NL Barry Bonds SFN 1381 186 204
2001 NL Barry Bonds SFN 1379 182 200
1920 AL Babe Ruth NYA 1379 189 182
1923 AL Babe Ruth NYA 1309 178 175
1921 AL Babe Ruth NYA 1359 179 175
1941 AL Ted Williams BOS 1287 177 174
2003 NL Barry Bonds SFN 1278 171 173
1927 AL Babe Ruth NYA 1258 168 173
1926 AL Babe Ruth NYA 1253 170 173
1931 AL Babe Ruth NYA 1195 162 172
Since Williams is hurt so much by Fenway Park he almost slips off the list entirely with only his 1941 season remaining. Ruth, however, now adds his 1927 and 1931 seasons when Yankee Stadium held a slight advantage for the pitcher.
Incidentally, Willie McCovey in 1969 moves up to tied for 11th with 165 when considering the tough hitting environment of Candlestick Park while Yaz in 1967 moves down to tied for 102nd at 148. So who is hurt most by taking into account park effects? As you might have guessed it is those who have played for the Colorado Rockies. In fact, Rockies take the top 35 spots when calculating the difference between NOPS and NOPS/PF with Todd Helton’s 2000 season taking top honors when his NOPS was 150 and NOPS/PF was 115. Conversely, Barry Bonds 2001 and 2002 seasons are most helped when park is taken into account raising his score by 18. For Cubs fans like me it’s interesting to note that Sammy Sosa’s 2000 season was tied for 2nd with a 15 point bump up to 149 once park effects were taken into account. Those who follow the Cubs know that weather patterns are the largest variable in whether or not Wrigley Field is a hitter’s delight or a pitcher’s best friend. Those who aren’t Bonds fans might take issue with assuming that the Giants home park hurts Bonds since it was built with a short right field porch with Bonds specifically in mind. Certainly Bonds, being a left-handed hitter, is hurt less by the park than are right handers and so I have a degree of sympathy for that argument. However, I don’t have any data that supports or contradicts the argument at this point. A similar argument could made against Ruth.
So does any of this change our perceptions of who had the greatest single seasons in history? Not really. Bonds, Ruth, and Williams still dominate the top spots and by virtue of Ruth taking 6 of the 10 a strong argument can be made that he was indeed the greatest hitter of them all.
On the other end of the spectrum Niefi Perez has somehow managed to grab two of the worst nine seasons in history with NOPS/PFs of 64 in 2002 and 71 in 1999.
Final Thoughts
Three additional thoughts might come to your mind when considering whether these were the greatest seasons in baseball history.
* Where's the defense? This ranking does not include defense and so can only be used as a ranking of the greatest offensive seasons in history. Although sabermetricians have tried for many years to develop defensive measures that quantify how many runs an individual saves for his team, in the end most of these schemes have difficulty. This is primarily because defense is a much more complex concept in baseball (more akin to defensive backs in football) than offense and doesn’t lead itself to quantification very easily. As Branch Rickey once famously said “There is nothing on earth anybody can do with fielding.” That said there are sabermetric measures such as Defensive Efficiency Rating (DER) and Zone Rating (ZR) that attempt to measure defense more accurately than the traditional counting stats that include put outs, assists, and errors. Bill James, in his book Win Shares, also tries to assign value to defense through a more holistic approach that takes into consideration run prevention at the team level.
* What about opportunities? As mentioned in the previous post one of the strengths of OPS is its simplicity. One of the costs of that simplicity is that OPS has nothing to say about the opportunity a player had to garner his OPS. In other words, which player is more valuable, one with an OPS of 850 who had 600 plate appearances or one with the same OPS who had 200 plate appearances? Obviously, the former since an 850 OPS is pretty good and so finding four players with an 850 OPS over 200 at bats will likely be difficult. In the rankings presented in this post this problem is largely ignored by selecting only those players with 502 or more plate appearances in a season, in other words by only selecting those players who played every day. To address this problem Runs Created per Game (RC/G) takes into account opportunities by considering how many outs a player has consumed – the most valuable resource a team has – while amassing their offensive numbers.
* What about Ty Cobb? Many readers will have noted that Ty Cobb is conspicuously absent from this list and that Cobb is often talked about in the same context with Ruth and Williams. In fact, Cobb first appears on the list tied for 33rd at 160 for his 1917 season right behind Sosa’s 161 in 2001. There are two reasons why this is the case. First, some of Cobb’s perceived value was his foot speed and base stealing ability, neither of which are particularly visible in OPS. Second, OPS is largely a measure of extra-base hitting and Cobb only hit as many as 12 homeruns twice. In his 1917 season, however, he hit 44 doubles, 24 triples, and 6 homeruns. The fact that OPS is correlated so strongly with run scoring indicates that players like Cobb who focused on hitting for average at the expense of power (assuming they could do either of course) did and continue to do a disservice to their teams by forsaking power. In short, if the often told story is true of Cobb hitting three homeruns in a game only to prove to writers that power hitting was not that difficult, then Cobb was mistaken in going back to his former style.
Finally, let's recalculate the 2003 leaders by applying both the correction for league and for park.
Raw OPS LgOPS NOPS/PF
Barry Bonds SFN 1278 749 173
Albert Pujols SLN 1106 749 154
Gary Sheffield ATL 1023 749 141
Jim Edmonds SLN 1002 749 140
Jim Thome PHI 958 749 135
Todd Helton COL 1088 749 129
Jason Giambi NYA 939 761 128
Carlos Delgado TOR 1019 761 128
Chipper Jones ATL 920 749 127
Manny Ramirez BOS 1014 761 127
It should be noted that the historical rankings take into consideration seasons since 1900 for players with 502 or more plate appearances and that for seasons with HBP and SF recorded they were taken into account. In addition, others have calculated similar adjusted values, some much more complicated, including the Adjusted OPS or OPS+ on baseball-reference.com and PRO+ in Total Baseball.
Posted by
Dan Agonistes
at
11:08 PM
2
comments
125x125_10off+copy.jpg)
