FREE hit counter and Internet traffic statistics from freestats.com

Monday, November 15, 2004

Baserunning 1992

Over the weekend I ran my baserunning framework for the 1992 season. In short, there wasn't alot that was surprising. The leaders in IBP were:


Opp Bases EB IB IBP OA IBR
Mark Whiten 43 76 62.68 13.32 1.21 0 4.40
Greg Gagne 40 70 58.13 11.87 1.20 0 3.92
Kenny Lofton 50 88 74.40 13.60 1.18 0 4.49
Bob Zupcic 25 45 38.27 6.73 1.18 0 2.22
Thomas Howard 25 42 35.97 6.03 1.17 0 1.99
Chad Curtis 28 50 42.96 7.04 1.16 0 2.32
Reggie Sanders 25 43 37.06 5.94 1.16 0 1.96
Rafael Palmeiro 48 85 73.39 11.61 1.16 0 3.83
Steve Sax 34 60 52.04 7.96 1.15 0 2.63
Kelly Gruber 22 36 31.25 4.75 1.15 1 1.48

Ok, so probably the most suprising things are that Rafael Palmeiro is in the top 10 and that Matt Williams comes in 21st out of 219 qualifiers. However, it must be remembered that Palmeiro was still just 27 years old in 1992 and that in 1993 he stole 22 bases and was caught only 3 times. His .85 IBP in 2003 puts him 323rd.

Other players in the top 20 are no surprise as well and include Lance Johnson, Henry Cotto, Dion James, Willie McGree and Bip Roberts while the bottom 20 include Kal Daniels, Fred McGriff, Kirt Manwaring, Kevin Mitchell, Tino Martinez, Kevin McReynolds, Charlie Hayes, and Mo Vaugn. Cecil Fielder comes in 203rd. Gary Gaetti takes the top spot for the most times thrown out advancing at 4. A top IBR of 4.49 for Kenny Lofton compares favorably with the 2004 data

I also noticed that the sheer number of opportunities for players is higher in 2003, hopefully not a result a miscalculation on my part. In 2003 for players with more than 20 opportunities the average was 34.9 while in 1992 it was 46.2. Part of this can be attributed to the increased offensive context of 2003 when 9.46 runs were scored per game versus 8.23 in 1992. In 1992 players did not get on base as often and the players behind them did not get hits as often to move them along. Interestingly, this means that from an absolute perspective a good baserunner in 2003 was worth more than a good baserunner in 1992. However, since each run was worth more in 1992 a good baserunner could more easily impact any particular inning or game.

Of course it's difficult to compare the same player across these two data sets to check for consistency since they occur twelve years apart which cuts down on the overlap and because you'd expect the IBP to decline during that time as the players age. That said, I did a quick comparison of 13 players that were in both studies. Of the 13, nine had a higher IBP in 1992 than in 2003 as you might expect. Overall the group's IBP was 1.04 in 1992 and 1.03 in 2003. Here are the results:

Opp Bases EB IB IBP OA IBR
Ken Griffey Jr.
2003 26 45 39.77 5.23 1.13 0 1.72
1992 33 54 47.87 6.13 1.13 0 2.02
Reggie Sanders
2003 35 55 50.88 4.12 1.08 0 1.36
1992 25 43 37.11 5.89 1.16 0 1.95
Kenny Lofton
2003 66 110 100.71 9.29 1.09 2 2.89
1992 50 88 74.48 13.52 1.18 0 4.46
Barry Bonds
2003 80 115 120.85 -5.85 0.95 2 -2.11
1992 55 84 86.34 -2.34 0.97 1 -0.86
Benito Santiago
2003 43 67 67.31 -0.31 1.00 2 -0.28
1992 21 28 30.54 -2.54 0.92 1 -0.93
Frank Thomas
2003 73 100 107.97 -7.97 0.93 0 -2.63
1992 70 111 106.37 4.63 1.04 0 1.53
Jeff Kent
2003 57 88 87.08 0.92 1.01 1 0.21
1992 50 76 71.42 4.58 1.06 0 1.51
Barry Larkin
2003 32 59 52.40 6.60 1.13 0 2.18
1992 39 61 58.01 2.99 1.05 0 0.99
Larry Walker
2003 54 96 86.56 9.44 1.11 0 3.11
1992 42 74 71.12 2.88 1.04 1 0.86
Gary Sheffield
2003 84 130 130.11 -0.11 1.00 2 -0.21
1992 34 56 53.45 2.55 1.05 0 0.84
Jeff Bagwell
2003 73 113 115.14 -2.14 0.98 2 -0.89
1992 44 69 69.07 -0.07 1.00 0 -0.02
Craig Biggio
2003 79 122 115.76 6.24 1.05 1 1.97
1992 67 94 100.75 -6.75 0.93 2 -2.41
Ivan Rodriguez
2003 61 92 90.80 1.20 1.01 1 0.31
1992 22 34 31.87 2.13 1.07 0 0.70

2003 763 1192 1165.36 26.64 1.02 13 7.62
1992 552 872 838.40 33.60 1.04 5 10.64

And from a team perspective they break down as follows:

Opp Bases EB IB IBP OA IBR
CLE 409 641 603.48 37.52 1.06 9 11.57
CAL 344 545 513.88 31.12 1.06 9 9.46
CHA 412 653 628.06 24.94 1.04 9 7.42
SFN 350 538 523.21 14.79 1.03 5 4.43
MIN 476 738 718.94 19.06 1.03 10 5.39
SLN 397 604 591.75 12.25 1.02 4 3.68
TEX 397 619 607.04 11.96 1.02 7 3.32
BAL 418 644 632.06 11.94 1.02 8 3.22
NYA 414 639 628.58 10.42 1.02 7 2.81
MIL 429 658 647.67 10.33 1.02 7 2.78
TOR 375 559 550.79 8.21 1.01 6 2.17
OAK 428 641 640.54 0.46 1.00 13 -1.02
NYN 347 536 535.92 0.08 1.00 11 -0.96
PIT 402 610 618.99 -8.99 0.99 11 -3.96
CHN 374 561 569.38 -8.38 0.99 10 -3.66
KCA 369 563 573.27 -10.27 0.98 14 -4.65
MON 399 588 600.18 -12.18 0.98 13 -5.19
LAN 370 539 552.93 -13.93 0.97 8 -5.32
PHI 403 588 604.49 -16.49 0.97 8 -6.16
DET 425 620 638.60 -18.60 0.97 9 -6.95
ATL 382 560 577.17 -17.17 0.97 12 -6.75
CIN 393 581 598.88 -17.88 0.97 10 -6.80
HOU 380 555 576.86 -21.86 0.96 9 -8.02
BOS 425 616 640.75 -24.75 0.96 13 -9.34
SEA 403 577 603.99 -26.99 0.96 8 -9.63
SDN 329 440 479.59 -39.59 0.92 12 -14.14

And as I found in 2003 the difference between teams is on the order of 75 bases or around 25 runs or 2.5 wins per season.

Sunday, November 14, 2004

Organizing Domain Logic

There are many different ways of organizing domain logic in Business Components that encapsulate calculations, validations, and other logic that drives the central functionality of an application or service. One of the books I like as a software architect is Patterns of Enterprise Application Architecture by Martin Fowler. In that book Fowler defines three architectural patterns (here listed in increasing order of complexity) that designers use to organize domain logic. I’ll explore each of these patterns as they apply to .NET development in a series of upcoming articles.

  • Transaction Script. This pattern involves creating methods in one or a few Business Components (classes) that map directly to the functionality that the application requires. The body of each method then executes the logic, often starting a transaction at the beginning and committing it at the end (hence the name). This technique is often the most intuitive but is not as flexible as other techniques and does not lead to code reuse. This pattern tends to view the application as a series of transactions.
  • Table Module. This pattern involves creating a Business Component for each major entity used in the system or service and creating methods to perform logic on one or more records of data. This pattern takes advantage of the ADO.NET DataSet and is a good mid-way point between the other two. This pattern tends to view the application as sets of tabular data.
  • Domain Model. Like the Table Module, this pattern involves creating objects to represent each of the entities in a system, however, here each object represents one Business Entity rather than only having one object for all entities. Here also, the logic for any particular operation is split across the objects rather than being contained in a single method. This is a more fully object-oriented approach and relies on creating custom classes. This pattern tends to view the application as a set of interrelated objects.

In additon to these patterns I'll also look at how the Service Layer pattern interacts with these. Should be fun.


Friday, November 12, 2004

Inerrancy and Ephesians 4:8 and Pslam 68

Recently I read both The Discarded Image (1964) and Reflections on the Pslams (1958) by C.S. Lewis for the first time. Both are wonderful books written in Lewis' later years (he died the same day Kennedy was assassinated) and are addressed to quite different audiences.

The Discarded Image is an introduction to medieval and renaissance literature born out of the lectures that Lewis gave after becoming the first Professor of Medieval and Renaissance Literature at Cambridge in 1953. In this short book Lewis explicates for his readers what he calls "The Model". The Model is no less than the lens in which people (probably more of the elite than the common man) in the medieval world viewed life. Lewis is often praised as one of the few people who really understood the medieval mind and this book makes you believe it.

Lewis begins with a short reconnaissance through important authors of both the classical and seminal periods (directly before the medieval period). And here his description of the works of Boethius (480-524) I found the most fascinating. If you've read much of Lewis you'll immediately see how influential Boethius was in Lewis' writings. Particularly, Boethius discusses the concepts of determinism and free will and Lewis sides with Boethius where he makes the distinction between God being eternal but not perpetual. God is outside of time and so never foresees, he simply sees. And for this reason he does not remember your acts of yesterday, nor does he forsee your acts of tomorrow. He simply experiences it in an Eternal Now (as the demon Screwtape says in Lewis' book The Screwtape Letters).

"I am none the less free to act as I choose in the future because God, in that future (His present) watches me acting."

Lewis then moves on to discussing The Model proper by breaking it into sections that include the heavens and the Primum Mobile (the outer sphere that God causes to rotate and that in turn causes the other fixed spheres to rotate), angels in all their hierarchical glory, the Longaevi (long-lived creatures like fairies, nymphs, etc.), the earth and animals, the soul, the body, history and the human past, and finally how these were taught using the seven liberal arts.

Although a more learned person would get more out of this book than I did, I can say it definitely opened my eyes and gives you just a glimpse of how a person in that time thought. I say a glimpse because we're so set in our modern ways of thinking that it is difficult even for a second to take in the night sky (as Lewis recommends) and imagine the vitality and purpose they saw in it.

The Reflections on the Psalms is a collection of essays on the Psalms that address issues and interpretations that Lewis came in contact with over the years. Unlike The Discarded Image, this book is not a scholarly work but as Lewis says in the first chapter, "I write for the unlearned about things in which I am unlearned myself".

For me, perhaps the most interesting part of this book is his discussion of "the cursings Psalms" such as Psalm 109 where the Psalmist rails against the "wicked and deceitful man" hoping all kinds of calamity for him and Psalm 137 where the poet says:

"O daughter of Babylon, you devastated one,
How blessed will be the one who repays you
With the recompense with which you have repaid us.
How blessed will be the one who seizes and dashes your little ones
Against the rock. "


Unlike the modern critic Lewis says that it is "monstrously simple-minded to read the cursings in the Psalms with no feeling except one of horror at the uncharity of the poets." Instead, Lewis looks past the sentiments that he says "are indeed devilish" and uses them as a spur to reflect on his own thoughts of uncharity and to look at the consequences of our own evil behavior on others. He also uses this occasion to see a general moral rule that "the higher, the more in danger". And as is typical of Lewis this comes from an unexpected direction with the idea that the higher or more developed moral code of the Jews made it more likely they would be tempted to a self righteousness contrary to those with lower morals. This too should be a lessen Christians can take from these Psalms.

The title of this post, however, is reflective of what Lewis later says regarding the both the inspiration of Scripture and "second meanings" in the Pslams. In the chapter titled "Scripture" Lewis says:

"I have been suspected of being what is called a Fundamentalist. That is because I never regard any narrative as unhistorical simply on the ground that it includes the miraculous. Some people find the miraculous so hard to believe that they cannot imagine any other reason for my acceptance of it other than a prior belief that every sentence of the Old Testament has historical or scientific truth. But this I do not hold, any more than St. Jerome did when he said that Moses described Creation 'after the manner of a popular poet' (as we should say, mythically) or than Calvin did when he doubted whether the story of Job were history or fiction."

In other words, Lewis did not hold to a belief in the inspiration of Scripture inclusive of the concept of inerrancy. To modern evangelicals a belief in inerrancy is indeed fundamental and usually is taken to mean that the "original autographs" of the Old and New Testaments were without error. Further, the concept of error encompasses not only where moral or spiritual matters are concerned but also scientific, historical, and literary. And of course this is why some in the evangelical community view Lewis' writings as dangerous.

While mulling over Lewis' view I then came upon his chapter on "Second Meanings in the Psalms" and to Paul's quote of Pslam 68 in Ephesians 4:7-8:

"But to each one of us grace was given according to the measure of Christ's gift. Therefore it says, 'WHEN HE ASCENDED ON HIGH, HE LED CAPTIVE A HOST OF CAPTIVES, AND HE GAVE GIFTS TO MEN.' " (NASB)

Paul is here "speaking of the gifts of the Spirit (4-7) and stressing the fact that they come after the Ascension". Unfortunately, neither the Greek nor the Hebrew Old Testament supports the reading "gave gifts to men" and instead say "received gifts from men" and in context refers to Yahweh and the armies of Israel as his agents, taking prisoners and booty (gifts) from their enemies (men) as Lewis points out. It appears that Paul is here relying on the Aramaic Targum, a Jewish commentary of the OT. Although most evangelicals assume that Paul is simply expanding the meaning of "received" to include the concept of giving, that seems a stretch to me. Instead it seems more realistic to assume that Paul used an incorrect translation.

For Lewis these are not problems because he viewed all Scripture as "profitable" when read in the right light and with proper instruction (human and through the Holy Spirit) per the teaching of 2 Timothy 3:16. Although I was once firmly in the evangelical camp on this issue, it now seems to me that this is the most reasonable way to understand the Bible and helps us to keep from getting bogged down in issues about possible or probable contradictions.

Greinke and Fergie

               BFP  K  BB  GB  OF IF  LD

Average 17% 10% 32% 22% 4% 13%

Anderson B. KC 745 9% 7% 28% 30% 4% 19%
Bautista D. KC 127 14% 10% 35% 24% 2% 14%
Greinke Z. KC 599 17% 6% 26% 29% 5% 16%
Gobble J. KC 638 8% 7% 31% 32% 6% 15%
May D. KC 832 14% 7% 27% 32% 4% 14%

I thought this was an interesting set of statistics I found in The Hardball Times 2004 Baseball Annual. It breaks down the outcomes by pitcher. Notice that all of them can be considered fly ball pitchers and Gobble and May particularly so. What doesn't bode well for Greinke is his line drive % sitting above average. I would think a pitcher's ability to suppress line drives would be directly related to his "stuff". One of the knocks against Greinke brought up by Bill James as I discussed in a previous post is that he doesn't have that pitch that can fool hitters. With that he still maintained an average strikeout rate because he locates so well.

In order to be successful a pitcher needs to adopt one of a few winning strategies. In the past I've speculated that these are:

1. Strikeout alot of batters in order to minimize the number of balls put into play and therefore balls that will be hits (Nolan Ryan)
2. Walk very few batters and give up very few homeruns to minimize the effect of the hits you do give up (Greg Maddux)
3. Walk fewer batters than average but strikeout more than average to minimize base runners and balls hit into play (Fergie Jenkins)
4. Rely on deception to decrease the number of hard hit balls thereby decreasing the pct of balls put into play that turn into hits (Charlie Hough)
5. Walk very few batters but rely on keeping hitters off balance to minimize base runners and minimize the number of line drives (Jamie Moyer)

To me Greinke fits into the mold of number 3 while Anderson, May and Gobble don't fit into any of the above. It's too early to tell for Bautista obviously.

Thursday, November 11, 2004

Royals Sign Truby

True to his word Allard Baird signed a "stop gap third baseman" today in soon to be 31 year-old Chris Truby. Truby has bounced around a bit playing with Houston, Montreal, Detroit, and Tampa Bay since 2000. The discouraging thing is that in 884 career plate appearances he has walked a grand total of 38 times. His lack of patience makes Terrance Long look like a "plate discipline guy".

To his credit he did have a good season in AAA hitting .300 with 25 homeruns and 41 doubles in 130 games in Nashville and was the Sounds MVP. He walked 47 times and struck out 96 times in his 466 at bats. By all accounts he is a good defensive third baseman and even played some second and short for the Sounds in 2004.

To me this move is basically a lateral one. Jed Hansen looks to be the same player with a little more plate discipline. I'd still rather see some kind of radical experiment with Ken Harvey until Mark Teahan is ready. Oh well.

Tippett on Speed

Although I knew Tom Tippett had some good research on his site I didn't know until today that he did a bit of analysis on the baserunning question I wrote about the other day. In his article Measuring the Impact of Speed he says the following regarding Ichiro Suzuki's 2001 season:

"The most surprising thing about this type of analysis is the relatively small number of baserunning opportunities we end up with. Players only reach base so many times in a season, and after you subtract the times when (a) the inning ends without any more hits, (b) they can jog home on a double, triple or homerun, and (c) they are blocked by another runner, players rarely get more than 50 opportunities per season to take an extra base on a hit.

Last year, Ichiro had 45 such opportunities, and he took 6 more bases than the average runner. He wasn't once thrown out trying in those situations. Six extra bases may not seem like a lot, but it was enough to qualify for our top baserunning rating."


That matches up very well with what I found. Of course, in my analysis runners get more opportunities because I don't take into account the runners in front of them and do give them credit for bases that they "should" get. So in 2003 Ichiro had 84 opportunities and was expected to garner 121.74 bases. He actually advanced 125 bases for an IBP of 1.03 and an IR of 1.08.

Tom also goes on to talk about other kinds of advancements and all the ways in which runners might be credited with extra bases. In Ichiro's case he finds that he took 30 extra bases which might equate to 6-12 extra runs using measures that are not normally accounted for.

Royals and K's

Ron Hostetter alerted me to Joe Posnanski's article on the Royals lack of strikeout pitchers. He also mentions that this should affect the long term evaluation of Jimmy Gobble as I articulated in a post last August. There I found that of young pitchers of comparable ability, the higher strikeout rate pitchers had careers almost twice as long and won more than twice as many games and that pitchers with long careers generally have above average strikeout rates when they're young.

In addition to Gobble, Miguel Asencio (4.45/9) and Kyle Snyder (4.11/9) are other suspects the Royals should look at when it comes to low strikeout rates. Joe didn't mention, however, that Tankersley's strikeout rate in AAA actually dropped last season more than two strikeouts per game while still managing to strike out 29 in 35 innings for the Padres.

From the article....Royals Strikeout Leaders since 2000

2000: Mac Suzuki, 135 (ranked 50th in majors)
2001: Jeff Suppan 120 (ranked 69th)
2002: Paul Byrd, 129 (ranked 53rd)
2003: Darrell May, 115 (ranked 72nd)
2004: Darrell May, 120 (ranked 67th)

Guidelines for Software Architecture

Recently, I've been thinking more about what a software architect does and trying to formulate some basic axioms or guidelines to keep in the front of my mind. To that end I thought Steve Cohen's thoughts on what "Enterprise Ready" means are useful.

Anyway, here are a few guidelines (not original of course) that I came up with for a course I teach on patterns and architecture in .NET that apply to a typical layered architecture (presentation, business, data).

  • Design components of a particular type to be consistent. By this we mean that components within a particular layer should use common semantics and communicate in a consistent fashion. One example is that in the Data Services layer all the Data Access Logic components should use the same mechanism (DataSets, data readers, custom objects) to marshal data.
  • Always run in process if possible. The basic facts of physics dictates that code running in-process is orders of magnitude faster than code running out of process. The reasons to run out of process include a need for process isolation, platform interoperability, and occasionally CPU overload (although “scaling out” often addresses this issue).
  • Only have dependencies down the call stack. In layered architecture it is important to not tightly couple the components lower in the stack to those higher. This allows the application to be modified more easily. Note that a strict rule such as never having the presentation call the data services directly is not implied by this guideline.
  • Separate model from view. Where possible try and abstract the data (the model) from the how the data is presented (the view). This can be done through the use of web user controls and the Bridge Pattern.
  • Use service accounts when possible. The use of service accounts to run server processes such as ASP.NET and Component Services allows for simplified security by not requiring delegation. It also allows connection pooling to occur.
  • Force common functionality up the inheritance hierarchy. One of the common patterns used is the Layer Supertype. This pattern promotes code reuse by implementing common code for an entire layer in an abstract base class whenever possible.
  • Keep policy code abstracted from application code. When possible, try to keep security, caching, logging and other “policy” code independent of the application or domain logic. This allows it to more easily be changed without affecting the functionality of the application.


Bingle Redux

Recently, I mentioned that in reading F.C. Lane's Batting I ran into the term "bingle". After having asked the eminent members of SABR about the origin of the term I posted the opinion that it was a contraction of the term "bunt single" and meant a slap hit, that being a style much in vogue in the deadball era. The term then became synonymous with "single" before dyeing out sometime in the 1950s.

Since then several more SABR members have chimed in. One thought that it was perhaps a blend of "bang" or "bing" and "single". However, now from Skip McAfee we have the following citations that seem to indicate that bingle was originally synonymous with single and perhaps later was used to refer to slap hits.

"Bingle is synonymous with 'base hit'. A player who bingles [note use as a verb] swats the ball safely to some part of the field where a biped in white flannel knickerbockers is not roaming at the immediate time" - Sporting Life, Dec. 15, 1900

"The big fellow grabbed three bingles in the afternoon contest, one of which was a smash over the fence that netted him a home run" - San Francisco Bulletin ,May 26, 1913

"Jack [Killilay] has dispensed eight bases on balls, hit one batter and permitted eight bingles, three of which were triples, in his stay on the mound" - San Francisco Bulletin, May 31, 1913

"In the third inning yesterday the first local batter up drew a bingle." - Youngstown Vindicator, July 29, 1898

And finally, to support the idea that bingle was used only later to refer to slap hits Walter K. Putney in Baseball Stories, the Spring 1952 issue noted that the term was "formerly the name for any kind of a hit", restricted the usage to a single of the "dinky" kind.

Also from 1913 Gerald Cohen's study of the San Francisco Bulletin noted the use of a) "bingle" as a verb (to hit or get a hit), b) "bingler" as a batter who gets a hit, and c) "bingling" as synonymous with "hitting".

Wednesday, November 10, 2004

Measuring Baserunning: A Framework

Who says there's an unemployment problem in this country? Just take the five percent unemployed and give them a baseball stat to follow.
--Outfielder Andy Van Slyke

In my previous two posts (here and here) I laid the groundwork for evaluating the baserunning of teams and players using play-by-play data from 2003. In the second post particularly, I showed the percentage of times players take the expected, +1, +2 number of bases in various situations and how often they get thrown out.

The Questions
Now I'm ready to make a first attempt at developing a baserunning framework in order to answer three related questions:

a) What player helped (or hurt) his team the most with his baserunning?
b) What team gained the most from smart or good baserunning in 2003?
c) What is the quantitative difference between good and bad baserunners?


Note that although this is my first attempt I'm putting this in public in order to get some feedback and certainly don't claim that this is the best method to use. I'm sure there are plenty of holes and problems, the two most pressing of which are that the sample sizes for a single season may not be large enough to differentiate ability from luck, and who hits after you has a large say in how many bases you advance. The former may be insurmountable with the limited data set I have although I'll try and correct for the latter as you'll see.

The Framework
The foundation for my baserunning framework is the table discussed in my previous post. You'll remember that it showed how often runners advance in various situations. For example, with a runner on first and nobody out when the batter singles to left field, the odds are:

Typ     +1    +2  OA

84.5% 14.1% 0.6% 0.7%

In other words, 84.5% of the time the batter stops at second, 14.1% of the time he advances to third, and .6% of the time he scores while .7% of the time he is thrown out on the bases. Using this set of percentages one can calculate the average number of bases advanced in this situation by multiplying the percentages by the bases gained. In this case (.841*1)+(.141*2)+(.06*3)-(.07*1) = 1.14. So when this event occurs a typical runner will advance 1.14 bases. Since this is the average across both leagues (I assumed it wouldn't be necessary to separate the leagues since there is significant overlap with interleague play, but more on that later) I call this Expected Bases (EB). The same calculation can then be done for the other 26 scenarios in the table (I did not use the Runner on 2nd Batter Doubles scenario in the calculations that follow since only one runner in all of 2003 was thrown out in that situation - the A's Mark Ellis). When this is done it turns out that the highest number of Expected Bases for any scenario is 1.86 which occurs with a runner on 2nd and 2 outs when the batter singles. The lowest number of Expected Bases is 1.14 for both the scenario given above and the same one but with 1 out.

It should be noted that for the total calculations below I also included singles fielded by other positions and so the actual number of scenarios is greater than 27. I found that shortstops and second baseman, for example, field a significant number of singles and to a lesser extent doubles, and that the typical number of bases advanced is similar to those fielded by outfielders. There were some plays were 0 was recorded as the fielder and so those were not considered.

As you probably anticipated one can then match up the baserunning situations for individual teams and players in order to compare the actual with the Expected Bases in each scenario. For example, Carlos Beltran of the Royals was at first base 9 times in 2003 when a batter singled to left field with 2 outs. In those situations he advanced to third twice and to second the other seven times. As a result he gained 11 bases. With those 9 opportunities he could have been expected to gain 10.39 (1.15*9) bases given the league average. As a result, he's credited with a positive .61 bases for this scenario, which I'm calling Incremental Bases (IB). When calculated for all of Beltran's opportunities in all opportunities we get a matrix like so where R1BD = Runner on 1st, Batter Doubles, Opp is the number of opportunities in each scenario, EB is Expected Bases, and IB is the Incremental Bases gained.

Sit Outs Fielded Opp Bases EB IB
R1BD 0 9 2 4 4.51 -.51
R1BD 1 7 4 11 8.95 2.04
R1BD 1 8 1 2 2.50 -.50
R1BD 2 8 2 6 5.45 .54
R1BS 0 3 1 1 1.02 -.02
R1BS 0 6 1 1 1.07 -.07
R1BS 0 7 1 2 1.13 .86
R1BS 0 8 1 1 1.28 -.28
R1BS 0 9 1 2 1.36 .63
R1BS 1 7 7 9 7.96 1.03
R1BS 1 8 4 5 5.11 -.11
R1BS 1 9 5 5 6.84 -1.8
R1BS 2 7 9 11 10.3 .60
R1BS 2 9 5 9 7.47 1.52
R2BS 0 7 2 3 2.72 .27
R2BS 0 8 1 2 1.63 .36
R2BS 0 9 4 6 5.62 .37
R2BS 1 3 3 3 3.58 -.58
R2BS 1 7 4 8 5.61 2.38
R2BS 1 8 3 6 4.90 1.09
R2BS 1 9 3 4 4.47 -.47
R2BS 2 4 2 2 2.05 -.05
R2BS 2 7 1 2 1.69 .30
R2BS 2 8 4 8 7.47 .52
R2BS 2 9 1 2 1.74 .25

So when all of these are summed we find that Beltran, in 72 opportunities gained 115 bases. He was expected to gain 106.7 so that puts him 8.36 bases gained above expected. What I like about this method is that it takes into consideration three context dependencies for the runner.

First, the handedness of the batters behind the baserunner are accounted for by looking at the fielder who fielded the hit. So if Mike Sweeney, a right handed hitter hits behind Carlos Beltran one would naturally assume that Beltran will have fewer opportunities to go from first to third because Sweeney is right handed. Beltran will not be punished in this system since we're comparing the number of bases he gained against the expected bases for the scenarios he was actually involved in. This system does not, however, control for how hard the batter hit the ball (which is possible given that there are codes in the data indicating line drive, fly ball, grounder) or park effects (Fenway Park might tend to decrease advancement to third on singles to left).

Second, this system takes into consideration the number of outs. This is important since we know from the table shown in the previous post that with two outs the probability of being able to advance extra bases often doubles. With this system Beltran does not get additional credit if he happens to be on base alot with 2 outs.

And most importantly, because each player will get a different number of opportunities both because of their own ability to get on base and because of the abilities of the batters following them, the sum of the bases gained can be divided by the Expected Bases to yield an Incremental Base Percentage of IBP. For Beltran that number is 1.08 and ranks him 56th among the 331 players with more than 20 opportunities in 2003, largely vindicating his reputation as an above average baserunner. In other words Beltran gained 8% more bases than would have been expected given his opportunities.

The Results
This calculation can then be run for all players and teams. The leaders in IBP for 2003 (more than 20 opportunities) are (you can find the complete Excel spreadsheet here):
	

Opp Bases EB IB IBP OA IBR
Miguel Olivo 26 43 35.45 7.55 1.21 0 2.49
Shane Halter 20 34 28.41 5.59 1.20 0 1.85
Chone Figgins 33 54 45.37 8.63 1.19 0 2.85
G. Matthews Jr. 50 89 75.75 13.25 1.17 0 4.37
Brian Roberts 63 104 89.37 14.63 1.16 0 4.83
Randy Winn 63 109 93.76 15.24 1.16 0 5.03
Denny Hocking 24 40 34.57 5.43 1.16 0 1.79
B. Phillips 31 53 45.85 7.15 1.16 1 2.27
Omar Vizquel 31 54 46.74 7.26 1.16 0 2.39
Rey Sanchez 35 62 53.93 8.07 1.15 0 2.66

While the leaders in total Incremental Bases are:

Opp Bases EB IB IBP OA IBR
Raul Ibanez 76 129 113.24 15.76 1.14 0 5.20
Randy Winn 63 109 93.76 15.24 1.16 0 5.03
Brian Roberts 63 104 89.37 14.63 1.16 0 4.83
Marcus Giles 74 122 108.29 13.71 1.13 0 4.53
Orlando Cabrera 67 116 102.61 13.39 1.13 0 4.42
G. Matthews Jr. 50 89 75.75 13.25 1.17 0 4.37
Luis Castillo 92 148 135.55 12.45 1.09 0 4.11
Albert Pujols 68 117 105.54 11.46 1.11 1 3.69
Derek Jeter 65 112 100.61 11.39 1.11 1 3.67
Todd Helton 84 142 131.16 10.84 1.08 0 3.58
Melvin Mora 59 99 88.27 10.73 1.12 0 3.54

In perusing the leaders in IBP and IB (we'll get to IBR in a moment) you do get the impression that these measures makes sense. The leaders in both lists tend to be those players we think of as fast and/or good baserunners. Even Larry Walker, not a particularly fast man but often mentioned as a good baserunner comes in 33rd out of 331 while players perceived as bad baserunners, such as Moises Alou at 277th and Ken Harvey at 289th, or simply slow (Jon Olerud at 321st and Edgar Martinez at 312th are near the bottom. Although the leaders in IB also reflect more opportunities, they seem to be pretty indicative of good baserunners with Raul Ibanez and Randy Winn leading the list.

Of course, I say overall because a pair of catchers, Miguel Olivo and Ben Petrick are the IBP leaders. This can be explained, however, by the fact that they had 26 and 22 opportunities respectively - very near the cutoff - and in the case of Olivo, he scored twice from first base on singles with two outs. Petrick was simply more consistent overall and scored from second all eight times he was there when a batter singled. Neither one was thrown out. This also points out that perhaps 20 opportunities is too low a threshold.

Another interesting case is the Tigers Alex Sanchez, a speedy man who often bunts for hits and who stole 52 bases in 2003. His IBP is only .90 ranking him 286th. A quick look reveals that while he's fast, he also takes lots of chances and was thrown out eight times, the most in the league, in 84 opportunities.

So in answer to question (a) above we can say that Raul Ibanez helped his team the most from his baserunning although Chone Figgins, Randy Winn, Brian Roberts, and Marlon Anderson are all right up there.

On the other side of the coin Mark Bellhorn is at .66 IBP good for 331st place. Bellhorn's poor performance was highlighted during his time with the Cubs in 2003 by his getting thrown out three times in twelve opportunities and only garnering 6 bases out of an expected 17. Some of this may be attributed to "Waving" Wendell Kim as I'll discuss below. For Chicago his IB was -10.92 and his IBP .35. His bad baserunning continued to some degree with the Rockies where his IB was -1.11 and his IB .94.

From a team perspective the leaders in IBP and IB were:

Opp Bases EB IB IBP OA IBR
COL 572 912 877.75 34.25 1.04 11 10.31
BAL 635 972 944.07 27.93 1.03 9 8.41
OAK 572 902 880.76 21.24 1.02 13 5.84
ANA 597 909 888.03 20.97 1.02 9 6.11
ATL 659 1005 982.46 22.54 1.02 14 6.18
CLE 551 850 831.83 18.17 1.02 15 4.65
MIN 626 959 944.88 14.12 1.01 15 3.31
SDN 594 900 887.49 12.51 1.01 6 3.59
NYN 521 778 769.58 8.42 1.01 13 1.61
KCA 681 1015 1004.02 10.98 1.01 12 2.54

As you can see Colorado had the highest IB followed by Baltimore. On the other end the Cubs had an IBP of .95 and an IB of -42.56. Using this we can tentatively answer question (b) as including Colorado, Baltimore, Oakland, and Anaheim as good baserunning teams. For question (c) the difference appears to be on the order of 75 or so bases per season that a great baserunning team takes over a bad one. It would be interesting to compare the 2004 numbers to see if there is any trend here and if the Cubs were justified in firing Kim.

The issue that this immediately raises is how much of Mark Bellhorn's poor performance can be attributed to his third base coach and how much to himself? As I showed earlier his performance definitely improved with the Rockies as did that of Jose Hernandez who's IBP was .96 with the Cubs and 1.05 and 1.08 with the Rockies and Pirates respectively although he had only four opportunities with the Cubs, far too small to say anything. So the question of whether there is a team bias at work and how large it may be is unknown.

From the team numbers it also appears there may be a league bias. Nine of the bottom ten teams are from NL while seven of the top ten teams are from the AL. I'll have to rerun the numbers to see if the probabilities are significantly different between the AL and NL but my assumption was that pitchers, while poor hitters, would not be significantly poorer in their baserunning ability. This may be incorrect or it could be that NL third base coaches are much more cautious with pitchers on the bases or that they take more chances when pitchers are coming up. Or a combination of all three.

Next Steps
So where to go next? It seems to me that next logical step is to translate IB into a number of runs gained or lost by individuals and teams. Two possible ways to do this occur to me.

One way would be to assign weights to the outs and advancements and simply sum them. For example, in the linear weights formula an out costs approximately -.09 runs (see my post on Batting Runs for a discussion of why -.09 instead of -.25) while a base gained from an intentional walk is weighted at .33. Using these values one can calculate an IBR (Incremental Base Runs) and see that the Rockies gained 10.31 runs while the Cubs lost 15.84 runs.

Opp Bases EB IB IBP OA IBR
COL 572 912 877.75 34.25 1.04 11 10.31
BAL 635 972 944.07 27.93 1.03 9 8.41
ATL 659 1005 982.46 22.54 1.02 14 6.18
ANA 597 909 888.03 20.97 1.02 9 6.11
OAK 572 902 880.76 21.24 1.02 13 5.84
CLE 551 850 831.83 18.17 1.02 15 4.65
SDN 594 900 887.49 12.51 1.01 6 3.59
MIN 626 959 944.88 14.12 1.01 15 3.31
KCA 681 1015 1004.02 10.98 1.01 12 2.54
NYN 521 778 769.58 8.42 1.01 13 1.61
SLN 611 940 931.92 8.08 1.01 14 1.41
TEX 550 825 822.77 2.23 1.00 14 -0.52
CHA 562 865 868.97 -3.97 1.00 8 -2.03
NYA 638 960 960.81 -0.81 1.00 20 -2.07
DET 446 639 642.49 -3.49 0.99 11 -2.14
TOR 656 1004 1007.05 -3.05 1.00 13 -2.18
FLO 560 823 829.39 -6.39 0.99 15 -3.46
SEA 664 976 983.03 -7.03 0.99 13 -3.49
PIT 585 867 877.20 -10.20 0.99 13 -4.54
TBA 603 894 906.19 -12.19 0.99 18 -5.64
CIN 508 754 767.83 -13.83 0.98 14 -5.82
MON 570 853 868.26 -15.26 0.98 14 -6.29
BOS 667 1011 1027.14 -16.14 0.98 12 -6.41
HOU 610 922 937.93 -15.93 0.98 18 -6.88
SFN 579 846 864.84 -18.84 0.98 12 -7.30
LAN 514 757 778.87 -21.87 0.97 10 -8.12
ARI 580 856 880.05 -24.05 0.97 10 -8.84
PHI 626 912 947.43 -35.43 0.96 15 -13.04
MIL 524 750 786.43 -36.43 0.95 17 -13.55
CHN 537 766 808.56 -42.56 0.95 20 -15.84

From an individual perspective Raul Ibanez leads with 5.20 IBR while Geoff Jenkins is last with -6.48 (he had an IBP of .80 in 53 opportunities). Looking at the spread this analysis indicates that good baserunning teams pick up about a win per year (assuming a win is purchased at the cost of 10 or so runs) over average teams and somewhat less than three wins over poor baserunning teams while an individual may be responsible for somewhat less than an extra win with his baserunning.

A second technique that could be used is to look at the run expectancy value for each situation before and after the play and calculate the difference. To me this makes a good deal of sense since it will have the tendency to weight the outs more properly and give more credit for actually scoring a run than simply advancing. A weakness of IBR is that an out at second base is treated the same as an out at the plate. I haven't yet run those numbers but may do so in the future. There is an additional problem doing it this way, however. The presence runners on the bases ahead of the runner we're analyzing will change the run expectancy even though of course the baserunner in no way controls what happens to those runners.

What's Missing?
I'm sure as you've read this you've thought of several things that might be included. Here is what I've identified.

1) This framework only includes three basic situations (runner on first batter singles, runner on second batter singles, and runner on first batter doubles). The situations could be expanded by looking at advancement on groundballs (so called "productive outs").

2) To get a complete view of the baserunning of an individual scoring on sacrifice flies, pickoffs, advancing on sacrifice hits, stolen bases, and even defensive indifference should be taken into account.

3) While some of the context is here accounted for, much else is not. For example, what if four of the eight times Alex Sanchez was thrown out on the base paths he was the tying run with two outs in the bottom of the ninth? Is it reasonable to punish him as severely as a guy who gets thrown out at third base with his team down 3-0 in the third? Obviously not.

4) The framework makes no allowances for the base ahead of the runner being occupied. This particularly effects hitters who are intentionally walked alot like Barry Bonds. For Bonds second was occupied 30 of the 47 times (64%) a batter singled with him there against the league average of 29.2%. In these circumstances the runner will find it more difficult to take an extra base, which will artificially hold down his IB and IBP values. The reason I didn't exclude these situations was because it would have further reduced the number of opportunities but a good case can be made for doing so.

5) It's not clear to me how a team would use this information to make better decisions except at the extremes: telling Alex Sanchez to stop trying to take an extra base every time you're on, firing your third base coach if you're the Phillies or Cubs, and using Gary Matthews Jr. and Shane Halter as my first pinch runners. In other words, while all of this is interesting and provides some quantification of baserunning, it's not very actionable for most teams or players. I realize that wasn't one of my questions when I started this but really worthwhile research should lead to something actionable.

Conclusion
In summary I want to reiterate that this is a first pass at analyzing and quantifying baserunning and for many of you (as for me) I'm sure has raised more questions than answers. I'd appreciate your thoughts in any case.

Long Shot

As many of us thought might happen Allard Baird traded Darrell May this week. He shipped him to San Diego along with Ryan Bukvich in exchange for Terrance Long and Dennis Tankersely. Both Long and Tankersely share the attribute that they were once much more highly regarded than they are now. I can't say I'm at all excited about the deal from the Royals perspective. However, after May's performance last season where he set a team record by giving up 38 homeruns and 105 extra base hits and was becoming a bit of a clubhouse cancer I doubt that Baird could have done much better. I think this is a good deal for the Padres who needed a fifth starter and whose wide open spaces at Petco Park should help May, the quintessential flyball pitcher (.76 GO/AO ratio last season). Ryan Bukvich also still has some upside at 26 years old.



  • Terrance Long. Wisely, Baird still understands that he needs a "corner outfield run-production guy" but as it sits right now Long will compete with Abraham Nunez for the right field job most likely. While Long is an upgrade to Dee Brown (who wouldn't be) because he has more power and is a better defender and runner, at best he's still a fourth outfielder who lacks plate-discipline (3.72 pitches per plate appearance career) as evidenced by his paltry career OBP of .319. He also seems to ground into alot of double plays and doesn't hit left-handed pitchers. That said, last year (.295/.335/.420) was Long's best since his rookie season when he hit 18 homeruns (.288/.336/.452). His VORP in 2004 was 12.3 with an Equivalent Average (EqA) of .268, very much in the middle of the pack for left fielders in the NL. Unfortunately, Long's power numbers and OPS have steadily declined from 1999 until a little rebound last year, something Long blames on trying to hit too many homeruns. He'll be 29 when spring training opens and so he's a long shot to break out. He has more value in the NL because he's left handed and seemed to pinch hit well so I wouldn't be surprised if his career trajectory takes him back in that direction next season. His price tag of $4.7M is about $4M too much for his skills but the Royals did get some cash to compensate in exchange for May's $3.225M 2005 salary. I assume this is a one-year deal. Although the story on the Royals web site mentions his 7 for 18 performance in the 2001 ALDS, overall he's .238/.304/.476 in the postseason in 69 plate appearances.


  • Dennis Tankersely. The upside on Tankersely is that he'll be 26 when spring training opens where Baird says that he'll compete for a 5th starter or long reliever role. Last season the right-hander definitely improved in AAA with a 3.15 ERA with a 1.26 WHIP at Portland. The interesting thing is that he seems to have found some control at the expense of strikeouts. His K/9 and BB/9 ratios the previous two seasons at Portland were 8.86/4.32 and last season they were 6.45/2.78. This probably indicates that he developed a third pitch (he was a 2 pitch pitcher prior to 2004 according to Baseball Prospectus) so his past performance may not be indicative of how he'll perform at the ML level next season if makes the team. Right now the Royals project Zack Greinke, Runylves Hernandez, Kyle Snyder, Brian Anderson, Jimmy Gobble, Denny Bautista, and Miguel Asencio to all compete for starting spots.

In other Royals news Zack Greinke came in 4th in voting for the Rookie of the Year where the A's Bobby Crosby dominated. David DeJesus came in 6th. Greinke and DeJesus did win the Royals pitcher and player of the year however.


Tuesday, November 09, 2004

Batting by F.C. Lane

I mentioned in a previous post that I recently came across a copy of F.C. Lane's (1885-1984) 1925 book Batting: One Thousand Expert Opinions on Every Conceivable Angle of Batting Science. This book was originally published as a supplement to Baseball Magazine at the price of $1 and contains quotes from over 250 players, managers, umpires, and owners on the nature of hitting covering everything from choosing a bat to dealing with umpires to pulling the ball and the merits of "slugging". Lane pulled these quotations from pieces he wrote for Baseball Magazine starting in 1910 or 1911. He continued to write for the magazine through the December 1937 issue and is prominately featured in the book The Numbers Game by Alan Schwarz. The copy I have was published by SABR in 2001 coming in at 218 pages in the original typesetting and contains an index of the quotations.

If you'd like to read something written by Lane see "The Base on Balls: Why Should the Records Ignore This Powerful Factor in Brainy Baseball?" posted by Cyril Morong from 1917.

Here are just a few of the tidbits that caught my attention while perusing the book.

On baseball records:
"The records are the whole thing in baseball. You can't make them too important and you can't be too careful in how you handle them."
- John Heydler, former Secretary of the National League:

"His batting average, then, is the ball player's principal stock in trade and is valuable to him in a personal and business sense. Where the records fail to give him credit due, he has suffered a genuine loss...In 1915, which was not my best year, I made twenty-four homeruns, a modern record. I also made thirty one doubles. I hit for 117 extra bases. The next best man in the League hit for 76. I scored more runs than anybody else on the circuit and drove in more runs. In all the really effective work of the batter I should have led the League by a wide margin." - Cactus Cravath

Lane was a pioneer in pointing out the one-dimensional nature of batting average and used Caravath as a prime example. As documented Schwarz observed in 1916 that batting average was an inadequate way of measuring the contribution individual players by remarking,

"Would a system that placed nickels, dimes, quarters, 50-cent pieces on the same basis be much of a system whereby to compute a man's financial resources? And yet it is precisely such a loose, inaccurate system which obtains in baseball..."

Although slugging percentage was adopted by the National League in 1923 and the American League in 1946 it would take many more years before weighted formulas for run production (Lane produced one himself) would start to take hold.

On slugging versus place hitting:
"My theory is that the bigger the bat, the faster the ball will travel. It's really the weight of the bat that drives the ball. My bat weighs 52 ounces. Most bats weigh 36 to 40 ounces...The harder you grip the bat, the faster the ball will travel...When I swing to meet the baseball I follow it all the way around." - Babe Ruth

"The sluggers have wrecked baseball. They are a thorn in the side of every pitcher. You never know when you have won the game. A homerun dumped into the stands may rob you of a victory any time until the last man is out." - Red Faber

"I do not lay my long hitting to any unusual strength. I believe it is due rather to meeting the ball fair, with a quick snap motion that sends it straight and true out over the diamond...It doesn't follow, however, that a heavy bat is necessary. Some sluggers use heavy bats. Chief Meyers did so and so does Babe Ruth. Most batters will find that an extra heavy bat cuts down the speed of their swing more than enough to offset what the extra weight of the bat can accomplish...I would say that in general the chop hitter would best use a heavy bat and the slugger a light one." - Rogers Hornsby

The debate between using a heavy versus a light bat is one that's had a long life. Over the course of time hitters have tended to move towards smaller bats with thin handles to increase the whip-like motion and therefore the speed of the barrel through the hitting surface. This is backed up by some recent research I wrote about in February. Both Harry Heilmann and Zack Wheat are quoted by Lane as favoring the lighter bat while Ken Williams and Hack Miller favor Ruth's approach. Note also that Robert Adair in his book The Physics of Baseball makes the point that the impact of bat and ball lasts about 1/1000 of a second. Therefore a tight or lose grip makes virtually no difference, nor does releasing the top hand after impact. The impact happens so quickly that the signal does not get processed until long after the ball is gone.

"Ruth is more than a slugger, he is a homerun hitter. Fortunately for him, he began as a pitcher. A pitcher is not expected to hit. Therefore, he can follow his own system without managerial interference. Ruth made the most of this opportunity...I have tried to make myself a batter, which is something quite different. A batter is a man who can bunt, place his hits, beat out infield drives, and slug when the occasion demands it, but he doesn't slug all the time." - Ty Cobb

I thought this was an interesting observation and seems reasonable to me. Had Ruth not been a pitcher a coach along the way might have tried to change his style thereby diminshing his power. It's also interesting to notice that in Lane's book Ruth's ability to hit homeruns and others ability to follow his example is chalked up to adopting a particular style or "speciality" of hitting rather than any influence of different baseballs as is the common perception. To me, the style that Ruth popularized along with the fact that whiter baseballs were kept in play after Ray Chapman's death after being hit with a Carl Mays pitch (Lane has some interesting quotes about beanballs as well including one from Walter Johnson that is especially emotional), and the reluctance of owners to make rules that handicapped Ruth in light of the Black Sox scandal all served to usher in the new slugger's era.

"If I had set out to be a homerun hitter, I am confident in a good season I would have made between twenty and thirty homers...I would naturally have sacrificed place hitting, which, to my way of thinking, is the supreme pinnacle of batting art." - Ty Cobb

This point is emphasized in other quotes Lane records as well. Many of the players, and Lane himself, seem to be of the opinion that while slugging is effective, it isn't pretty and requires less skill. Almost to a man Lane records that everyone thinks Ruth the greatest slugger (and slugging as a specialty) but Cobb the greatest hitter (there are few votes for Honus Wagner as well). Sabermetrics of course, would argue that Cobb should have taken his 30 homeruns. Overall Ty Cobb and George Sisler are most appealed to on the subject of hitting while Ruth and Hornsby are close behind.

There are countless other fascinating little insights as well that any fan of baseball history will enjoy. Highly recommended.

Monday, November 08, 2004

Measuring Baserunning: Setting a Baseline

In the interests of furthering the state of sabermetric knowledge per my previous post I compiled the probabilities of runners advancing in various situations using play-by-play data from the 2003 season. The scenarios I looked at included:

  • Runner on 1st, Batter Singles
  • Runner on 1st, Batter Doubles
  • Runner on 2nd, Batter Singles
  • Runner on 2nd, Batter Doubles

For each situation I calculated the odds of advancing a "typical" number of bases (1 for a single, 2 for a double), the odds of going one better, the odds of going 2 better (where applicable), the odds of getting thrown out advancing, and the percentage of the time a runner occupied the base directly in front of the runner. I then create subtotals for each situation by the number of outs and the outfielder who fielded the hit. For example, for a runner on 1st when a batter singles the values are:



Outs To Typ +1 +2 OA Next Base Occ
All All 70.5% 27.2% 0.9% 1.4% 29.2%
0 All 73.4% 25.0% 0.5% 1.2% 21.1%
7 84.5% 14.1% 0.6% 0.7% 21.6%
8 68.6% 30.1% 0.3% 1.1% 25.0%
9 59.7% 38.3% 0.6% 1.4% 16.5%
1 All 72.4% 25.5% 0.7% 1.3% 30.6%
7 84.7% 13.4% 1.0% 0.9% 31.5%
8 70.3% 28.6% 0.4% 0.7% 34.0%
9 58.1% 39.0% 0.9% 2.0% 27.9%
2 All 66.3% 30.7% 1.4% 1.6% 33.9%
7 81.1% 15.8% 1.4% 1.6% 33.8%
8 60.0% 35.9% 1.8% 2.3% 33.1%
9 48.3% 49.7% 1.0% 1.1% 31.1%
All 7 83.4% 14.4% 1.0% 1.1% 29.6%
All 8 65.8% 31.9% 0.9% 1.4% 31.5%
All 9 55.3% 42.3% 0.8% 1.5% 25.8%

The value in breaking these numbers down can be illustrated by noticing how runners advance from first to third less than a third of the time (27.2%) overall but almost half the time (49.2%) with two outs when the ball is hit to right field. It's also interesting to notice that runners get thrown out most often in this situation with 2 outs when the ball is hit to center field and that runners only advance to third base 14.4% of the time when the ball is hit to left field.

Because I couldn't format the results as I wanted on this blog I've posted the complete results for all scenarios here.

More on how we can use these numbers to evaluate teams and players next time.

Win Expectancy and Post Season Awards

Nice piece in the NY Times by Alan Schwarz, author of The Numbers Game, on Wins Relative to Average Player (WRAP). This stat was created by Brian Lonergan, a young economist from Yale, and Ben Polak his former dissertation adviser.

Essentially, this technique uses win expectancy tables to calculate the odds of winning before and after each of the 186,000 plate appearances and assigning responsibility to players involved in the play. Each player's contributions are then summed and a total number of wins they are responsible for calculated. The sum of all contributions for all players on the team equal the number of games the team is above .500.

This technique borrows heavily from the Mills Brothers Player Win Averages (PWA) methodology created in the late 1960s as well as Bennett and Flueck's Player Game Percentage (PGP) documented in chapter 10 of Curve Ball. The difference is that this system takes the next step and translates the probabilities into wins ala Bill James' Win Shares system.

Their results for 2004 include:

"A.L. M.V.P. Sheffield, with a 5.60 WRAP, edges out the Angels slugger Vladimir Guerrero (4.45). Ramirez ranks sixth because more of his hits came in game situations that did not have a big effect on the outcome of Boston games.

N.L. M.V.P. Bonds (12.16) tramples the runner-up, Albert Pujols of the Cardinals (6.85), demonstrating that Bonds's walks did indeed help the Giants by setting the table for subsequent, however disappointing, hitters.

A.L. CY YOUNG Despite having a higher E.R.A., Schilling tops Santana, 5.15 to 5.01, because he often performed in hitter-friendly Fenway Park (yes, WRAP accounts for this) and because he thrived in particularly tight situations. But the two were bested by Twins closer Joe Nathan (5.47), whose 1.62 E.R.A. and 44 saves do not truly quantify how many games his late pitching helped decide. (WRAP leans toward relievers because, although they influence fewer at-bats, the at-bats are inherently more crucial.)

N.L. CY YOUNG Once again, closers dominate, with the Dodgers' Eric Gagne (5.69) beating the Astros' Brad Lidge (5.55). San Francisco's Jason Schmidt (5.39) wins among starters, well ahead of the more popular candidates for postseason awards, Randy Johnson of Arizona and Roger Clemens of Houston."

Win Shares of course equate to one-third of a run and so comparing the 2004 Win Shares as calculated on The Hardball Times you get the following for those mentioned in the article (their winners bolded). WS=Win Shares, WSAA=Win Shares Above Average, Wins=WSAA/3 for comparison with WRAP.



WS WSAA Wins
Sheffield 31 15 5
Guerrero 29 12 4
Ramirez 28 12 4
Santana 27 15 5
Schilling 22 9 3
Bonds 53 37 12.3
Pujols 40 21 7
Rolen 38 21 7
Beltre 37 19 6.3
Abreu 37 18 6
Gagne 16 8 2.7
Lidge 17 9 3
Schmidt 19 8 2.7
Clemens 20 10 3.3

In evaluating the two systems it seems to me that WRAP overvalues closers because of the highly volatile situations in which they work (the danger of using situation-dependent measures like this) while Win Shares undervalues pitchers (Clemens was the highest rated pitcher in the NL and came in 33rd in total Win Shares behind Miguel Cabrera, Derek Lee, and Luis Castillo). Note, however, that Bonds, Pujols, Sheffield, Guerrero, and Santana were credited with a very similar number of wins in both systems.

You can also note that the ranges of both systems are similar and so generally you can say that MVP-type players may be worth an additional 5 to 8 wins (except for Bonds).

The Argument From Reason

A couple weeks back I attended a lecture by Dr. Ned Keller, a physicist who teaches college in Grand Rapids Michigan. The lecture was entitled "The Nature of Knowledge" and was the first in a series he did discussing "Theories of Origins" at an area church. In it Dr. Keller provided a basic outline of how we know what we know in order to provide a foundation for what he was to talk about the rest of the weekend. Although I didn't get to attend any but the first lecture, I found it interesting that among his preliminary topics for the weekend was:

The preliminary question is whether the universe is best described with a natural or a supernatural worldview. What are miracles? Can they occur and have the occurred?

To answer this question Keller went to chapter 3 of C.S. Lewis’ 1947 book Miracles titled “The Cardinal Difficulty of Naturalism”. In that chapter Lewis makes an argument for God and against Naturalism (that nature is the “whole show”) based on the existence of human reason. Interestingly, after Miracles was first published a debate on Lewis’ argument was held at the Oxford Socratic Club between Lewis and philosopher Elizabeth Anscombe. By all accounts Lewis lost the debate and subsequently rewrote the chapter for subsequent editions of the book.

In his lecture Keller talks both about chapter 3 and about miracles in general.

Lewis’ argument for the existence of the supernatural is known as the “Argument From Reason” or the "Argument From Mind".[1] In short Lewis puts the argument in syllogistic terms like so:

  • If supernaturalism is not true, then Nature is all there is and so "every finite thing or event must be (in principal) explicable in terms of the Total System [i.e. Nature]." This is called physicalism by Moreland.
  • Therefore if any one thing can be shown to be not explicable in terms of the Total System, then naturalism is not true
  • Human reasoning can be shown to be something not explicable in purely physical terms and so naturalism is therefore not true. Man’s "rationality is the little tell-tale rift in Nature which shows that there is something beyond or behind her." [2]

Lewis then defends his conclusion by starting with the premise that all knowledge depends on the validity of our reasoning. If reasoning does not lead us to true conclusions about the world around us, but is rather a product of feelings in our own mind, then all science and all knowledge is worthless. This, says Lewis, points out that strict materialism or physicalism is self-defeating. In other words physicalism may be true, but one cannot argue that it should be believed based on evidence or reasoning.

Lewis then goes on to argue why it is that human reasoning cannot be explained in terms of the “whole show”. He illustrates this through the two different senses of the word "because". In the first sense because can be used to mean a cause and effect relationship ("Grandfather is ill today because he ate lobster yesterday"). In the second sense because is used in a Ground and Consequent relation, for example, "Grandfather must be ill today because he hasn’t got up yet (and we know he is an invariably early riser when he is well"). The first sense indicates a connection between a state of affairs while the second is a logical relation between beliefs that involves an act of knowing or seeing or rational insight. Another example of the second sense of because is the mathematical reasoning if A=B and B=C then A=C.

Lewis then explains that every event in nature, including our very thoughts, must be of the first type if naturalism is true. If this is the case, then when we ask "Why do you think this?", the actual answer must always begin with a Cause-Effect style because. As a result, all of the thoughts that go into answering the question lie in a cause-effect relationship to one another including the final answer. But we know that to be caused is not to proved and so the physicalist, if he is consistent, must admit he has no way to know whether what he thinks is true. He has no way to bridge the gap between the two distinct senses of because. In fact, as Lewis notes, in argumentation people often act as if the two were unrelated so that if a person can find some bit of background about you that might indicate why you believe something (Cause-Effect) they can more easily discount your position. However, in our experience we know that not all of our thoughts are based on wholly on Cause-Effects relationships (we don’t draw all of the inferences possible from each thought). Some of our thoughts can cause other thoughts by being seen to be a ground for them (Ground-Consequent). Therefore, since some of our thoughts can be shown to be true acts of knowing or seeing that cannot be accounted for by naturalism, then naturalism is false. This then explains why all human reasoning and therefore science and knowledge must be thrown out for the physicalist since it depends on Ground-Consequent style thinking including the physicalist’s own conclusion that nature is the whole show. This is why that position is self-refuting. As Lewis conludes:

"But this, as it seems to me, is what Natualism is bound to do. It offers what professes to be a full account of our mental behaviour; but this account, on inspection, leaves no room for the acts of knowing or insight on which the whole value of our thinking, as a means to truth, depends."

He also goes on to address the naturalists claim that our reasoning is the product of natural selection and/or cultural evolution. Against natural selection as the origin he argues that natural selection can only improve man’s physical responses to the world around him and could never in principle develop a relationship between knowledge and truth since there is no connection between the two, no way to bridge the gap. Against cultural evolution or the belief that over time, men were conditioned to make inferences based on experience (where there is smoke there is fire), Lewis argues that inferences are the basis of animal, not human reasoning. The real difference between animal and human reasoning is that human reason need not appeal to experience at all. For example, our belief that A=C as above is not derived from our experience that we’ve never not know A to equal C. Rather, it is based on a real insight that "it must be so". These are the insights that a physicalist cannot explain.

[1] There are several other arguments for God’s existence that can be used including the Ontological Argument, the Cosmological Argument (which has three different forms), and the Argument from Design. The books Scaling the Secular City by J.P. Moreland and Reasonable Faith by William Lane Craig have very readable introductions to these arguments.

[2] Moreland also gives other reasons for thinking that there are entities that cannot be explained naturalistically including moral values, numbers, and universals (concepts such as color)


Saturday, November 06, 2004

Neifi the Terrible

So the Cubs signed Neifi Perez to a one-year contract.

With tomorrow being Sunday let us all take a moment to pray that Jim Hendry and Dusty Baker don't actually think that this 32 year-old with a career .301 OBP and .380 slugging percentage (despite playing six of his nine season in Colorado) should be allowed to compete for a starting job.

Despite his 6 for 6 debut with the Cubs and .371 average in 62 at bats he still only managed .255/.296/.336 last season, a shade under replacement level (-.2 runs) based on the VORP numbers from Baseball Prospectus. He also has the distinction of capturing two of the worst nine offensive seasons in baseball history with his normalized OPS (adjusted for park factor) values of 64 in 2002 and 71 in 1999.

At best Perez, with his more than adequate defense, could take the utility infield spot of Jose Macias whose .268/.292/.376 performance was a touch better (+1.9 VORP) than "Neifi the Terrible".

Gruenwald USA Today Athlete of the Week

My sister-in-law's husband Jim Gruenwald was named USA Today athlete of the week. He battled back from a serious injury at the 2003 World Championships to wrestle in Athens for the US Greco-Roman team at 132 pounds. He also wrestled in Sydney in 2000, placing 6th.

Thursday, November 04, 2004

The Cubs and OBP

Interesting article about OBP and the Cubs focusing on Mark Bellhorn. It echoes what many of us have been saying for several years, the Cubs just don't get it when it comes to getting guys on base.

"The Cubs finished 11th in the NL in team on-base percentage. Despite a great starting-pitching rotation, they failed down the stretch largely because they didn't score runs, even against the bottom-feeding Mets and Reds. You can blame fatigue, but you also can blame bad approaches at the plate, including too much first-pitch swinging and swinging at bad pitches."

The author also dredges up Dusty's quote on walks "clogging" up the bases that I've referred to several times.

Big League Manager - Desktop Version

Since most folks don't have access to a Pocket PC I created a desktop version of the Big League Pocket Manager simply called Big League Manager. You can download it here. Simply double-click the .msi file to install. This first version has the same functionality as the PPC version but I'm looking for ways to expand it - for example by offering multiple hitter profiles based on lineup position perhaps.

Keep in mind that this is a quick and dirty first version so if you find problems feel free to contact me.


Wednesday, November 03, 2004

Actuaries and Sabermetrics

I was alerted to an article titled "Stat of the Art: The Actuarial Game of Baseball" published in a journal for actuaries, Contingencies May/June 2004. The article itself is just a recapitulation of Moneyball, Billy Beane, and Bill James with the Red Sox and is old news to anyone who's been following sabermetric developments. Two factual errors I noticed on skimming the article.

  1. It says Bill James is from "Lawrenceville" Kansas instead of Lawrence
  2. It says Bill James invented On-base percentage. A precursor to OBA was actually developed and then dropped as early as 1879 but OBA became more recognized in the 1950s.

What is more interesting are the eight proposals for "new" statistics from readers of Contingencies that Bill James reviews.

In a nutshell these are:

  • The Reliever Effectiveness Ratio by Damian Birnstihl. This is simply a recalculation of ERA by assigning a pitcher a half run for each runner he let on that scores and a half run for each runner he lets in after he comes in as a reliever. This is a simpler version of what Ari Kaplan did a decade ago as I blogged about here.
  • Rearranging the Starters by Aryeh Bak. This is a study that concludes that rearranging your starting pitchers based on the opponent's starters could gain a couple extra wins during the season. James points out that this question has been studied before by Dallas Adams and Tom Tippett and that there are so many assumptions you have to make that in practice this doesn't work. To me what is more important is getting your best pitchers on the mound more frequently (hence the four-man rotation), not working matchups.
  • A New Wrinkle by Rod Keefer. This is a "new' stat called "Run Production Index" (RPI) that takes runs produced (R+RBI-HR) and divides it by at bats. James points out that runs produced has been around since the 1950s and then goes into its flaws. Keefer also proposes calculating "bases advanced per at bat" (BAAB) or how many bases each offensive event resulted in. To me this makes some sense but is awfully dependent on the other hitters in the lineup.
  • The Three-Year-Ago Correlation by Paul Conlin. Conlin said that he found that team winning percentages were more highly correlated with results from three years ago than from any previous year. James responds with a study that refutes it and shows that the average change in wins from one year to the next is around 10 and then steadily increases each prior year.
  • Starting Pitchers and Relief Pitchers by Mark Seliber. A spreadsheet with the same stat as RPI as well as new stats for starters and relievers. Nothing new or interesting here.
  • The Efficiency Rating by Myron Kraynyk. This is a more interesting attempt that uses the same idea as BAAB but also considers the total possible bases available in each plate appearance. These are summed over all plate appearances and divided to yield an Efficiency Rating. No data is provided however and it suffers from being very context dependent.
  • The Base Advancement Percentage by Richard T. Newell Jr. This stat is essentially the same as the Efficiency Rating.
  • The Ultimate Baseball Statistic by Spencer M. Gluck. This is essentially another version of Player Win Averages (PWA) developed by the Mills brothers in the late 1960s. James isn't sold on these techniques and says that systems like this that take into account everything end up producing nothing of value since there are too many unknowns that have to be glossed over. Another way of saying it is that a good baseball analyst contributes to the discussion by answering specific questions that are actionable. In summary James says:

"Good analysis never begins with statistics. Good analysis always begins with a specific question, and a question which is of interest to baseball people, whether they are actuaries or artists or aging scouts. Proposing a system which instantly evaluates everything that every player does is analogous to fixing insurance rates for drivers by attaching a box of sensors to the hood of every automobile and keeping track of how often every driver does something dangerous, and calculating exactly how dangerous that was. It’s not the real world. It’s not practical, and it’s not useful. Maybe, in 50 years, it will be practical or it will be useful, but it’s not now. Our general knowledge is limited by our specific knowledge. Our ability to have an impact on the discussion cannot be larger than our ability to find a question which has an actual answer."

What also interested me about these "new" statistics are of course that they're not new and that even in this small sample four of the eight people had an idea that was the same or very similar to one of the others. More to the point, these are the same kinds of ideas that baseball analysts have had going back to F.C. Lane in the 1920s, George Lindsey in the 1950s, Earnshaw Cook, the Mills brothers, Pete Palmer and on and on. Once again this seems to call out for a place where this kind of information is collected.

Tuesday, November 02, 2004

Measuring Baserunning: A Preliminary Attempt

Another great quote from Rich's post yesterday on the 1984 Baseball Abstract was James' contention that Project Scoresheet was going to definitively answer questions about baserunning.

"Baserunning is perfectly measurable; it can be easily defined and, given properly maintained scoresheets, easily researched. Our lack of knowledge on the subject is attributable entirely to record-keeping decisions that were made a little over a century ago and have never been intelligently or systematically reviewed. We know so much about hitting that we can talk about it forever and measure it with extraordinary precision because a few men, at the beginning of Time, made some very good decisions about how to record and organize information, decisions that are now so natural a part of our thinking about the game that it is difficult even to see that any decision had ever to be made.

For this we applaud them. Their decisions about baserunning and fielding were much less wise. They failed to address many issues, and drew arbitrary lines where they drew them at all, and time has laid waste to their designs."

Rich then goes on to say:

"If this information is known today, it sure isn't widely disseminated. Why don't we know how often (in absolute terms and as a percentage of opportunities) various runners go from first to third on a single, first to home on a double, or second to home on a single? How often does Ichiro Suzuki reach base on an error as opposed to the average batter? Are we limited in recording the data or in distributing the data? Until this information is made available to the public, we will be limited in our ability to fully understand and appreciate all the nuances of the game and its players."

Well, in the interests of picking up the challenge I loaded the 2003 play-by-play data (found on the Yahoo group stats_software) into SQL Server last night and wrote a few quick queries to provide a baseline to answer a couple of Rich's questions and for future work.

Situation
Man on First - Batter Singles
Opportunities: 10430
ToThird: 2841 (27.2%)
Scores: 94 (0.9%)
Out Advancing: 147 (1.4%)

Man on First - Batter Doubles
Opportunities: 2979
Scores: 1315 (44.2%)
Out Advancing: 95 (3.2%)

Man on Second - Batter Singles
Opportunities: 6128
Scores: 3703 (60.4%)
Out Advancing: 217 (3.5%)

I was somewhat surprised how often a runner scores from first on a double (44.2%) and how few times a runner on first is thrown out advancing. Breaking these numbers down by the number of outs (which I plan to do) would also show a difference I assume.

From a team perspective the leaders in these categories were:

Man on First - Batter Singles
ToThird: Colorado and Minnesota (33.3%)
Scores: San Diego (2.0%)
Out Advancing: Houston (3.4%)

Man on First - Batter Doubles
Scores: Montreal (53.8%)
Out Advancing: Cubs (6.4%)

Man on Second - Batter Singles
Scores: Oakland (66.0%)
Out Advancing: Cubs (6.3%)
Cubbs fans should probably not be surprised that former third base coach "Waving" Wendell Kim in 2003 contributed to 20 runners being thrown out on the basepaths, tied for the most in baseball with the Yankees (the Padres had only 6).

These metrics get a bit difficult, however, when you consider the personnel the third base coach has to work with and judging what is and is not an opportunity In these numbers I've looked at all opportunities and not only those where the next base was unoccupied. There is an argument to be made that if the next base is occupied then the runner may be hindered in their attempt to take the extra base. Overall, the next base was unoccupied around 70% of the time so I went ahead and included all opportunities.

Overall, a quick way to measure a team's baserunning skills might be to see how many times they took at least one "extra" base. The percentage leaders are:

Colorado 46.6%
Oakland 46.1%
Minnesota 44.1%
Cleveland 44.0%
Baltimore 43.6%

The Cubs were last at 35.5%.

As far as individual leaders in these categories go:

Man on First - Batter Singles (more than 20 opp)
ToThird: Jose Guillen (59.1% 26/44)
Scores: Luis Matos (12.9% 4/31)
Out Advancing: Matt Lawton (11.1% 3/27)

Man on First - Batter Doubles (more than 10 opp)
Scores: Miguel Tejada (81.3% 13/16)
Out Advancing: Juan Encarnacion (18.2% 2/11)

Man on Second - Batter Singles (more than 20 opp)
Scores: Jose Guillen (90.9% 20/22)
Out Advancing: Tony Womack (14.3% 6/42)

Interestingly Juan Encarnacion was second in getting thrown out when on second with a single 13.6% of the time (3/22). So even from these small samples it appears that perhaps good and bad baserunners can be identified. Jose Guillen looks pretty good here while Juan Encarnacion does not.

My ultimate goal here is to assign weights to the various baserunning plays and come up with a single number that estimates how many runs a player gained or lost for his team on the bases. This methodology has some limitations and doesn't provide a complete picture, e.g. looking only at advancement in this manner does not include advancement on errors, does not measure the impact of holding the runner, doesn't take into account park effects (the Green Monster likely has a negative impact on the times a runner can go from first to third on a single), and as I mentioned there is a strong weighting based on the third base coach that would have to be separated from individual players. Also the sample sizes in a single season for individuals is really very small. I think you'd need 5 to 10 years worth of data to start to get something meaningful.

All of these limitations and more are what make analyzing baserunning difficult.

Monday, November 01, 2004

"THEODORE! with all thy faults — "

Excellent column by George Will that summarizes my reasons for voting for Bush, which in many ways is largely a vote against Kerry.

Baseball from the Outside

Rich Lardner over at Rich's Weekend Baseball BEAT had a great post today about Bill James' 1984 Baseball Abstract. In particular Lardner talks about the essay "Inside-Out Perspective" where James discusses the trend toward "inside stuff" in sportswriting. While that topic seems a bit dated, the most interesting aspect of the essay is his description of what he is doing in the abstracts:

"This is outside baseball. This is a book about what baseball looks like if you step back from it and study it immensely and minutely, but from a distance."

He then goes on to say:

"But perspective can be gained only when details are lost. A sense of the size of everything and the relationships between everything--this can never be put together from details. For the most essential fact of a forest is this: The forest itself is immensely larger than anything inside of it. That is why, of course, you can't see the forest for the trees; each detail, in proportion to its size and your proximity to it, obscures a thousand or a million other details."

I've read that book and essay several times over the years but it wasn't until attending almost three dozen Royals games this season and being immersed in the detail of each play while scoring for MLB.com that the essay brought home to me a new meaning.

I had assumed that viewing the game from that perspective would lead to insights that I hadn't previously been privy to. And indeed I did learn more about pitching patterns and, for example, the importance of establishing the strike zone early in the game. However, overall what I realized from personal experience was that it was very difficult to be immersed in the details of each game and yet maintain perspective on the bigger picture - a picture that looked at season and career trends as well as strategies. By the end of the season I found it quite difficult to separate in my mind events that happened in different games and found that I was thinking about players and had made judgments in terms of just a handful of the plays or situations I remembered clearly. And I only attended less than half the Royals home games.

I think this goes a long way towards explaining why sportswriters and those inside baseball have long resisted sabermetric analysis. They are simply too close to the game. There are too many discrete events in the course of a baseball season and too much randomness thrown in for human minds to sort it all out. Paradoxically, those who have witnessed every inning of every game think they must know more about the game since they've seen more. In actuality, the limitation of our brains makes us susceptible to forming invalid conclusions based on the limited number of events we remember and the emphasis we place on them. In some ways those who are closer actually know less. And so of course, when analysis from an outside perspective calls into question the existing paradigm, there will be fierce resistance. This problem only grows worse when you also consider the personality and relationship issues that are a part of any human endeavor and that go into decision making within a ballclub.

To me this means that a baseball analyst, in order to be effective, must purposefully take a step back from the trees in order to see the forest and work from the proper perspective. The kinds of questions that an analyst can help answer then are by nature big picture questions: Is it generally a good idea to bunt in a particular situation? Is developing a closer useful? Does drafting high-school pitchers pay off? Does guarding the lines help or hurt in the late innings? At what age do players generally start declining in productivity? These are the kinds of conclusions Paul DePodesta talked about with his "Be the House" mantra - general strategies that only pay off over the long haul.

This is precisely the kind of role it appears James now has with the Red Sox. He is not involved in the day to day operations of the team and instead creates studies from an outsider's perspective that GM Theo Epstein and the management team can use in their decision making process.

This also means that if news organizations and other teams want to benefit from that kind of analysis they would do well to hire someone from the outside and then keep that person at arm's length.