When Do Completion Percentage, Yards Per Attempt, and Adjusted Yards Per Attempt Stabilize for College Quarterbacks?

Day 1 of the 2018 NFL draft is shaping up to be QB-centric, with five passers projected to come off the board. Owing to this, the commentary that’s entered my social media orbit over the past month has been completion percentage-centric. As I’ve never paid much attention to the draft, I was surprised to have a human emotion upon consuming said commentary. It didn’t take long to figure out why: No one knows, statistically speaking, how much college completion percentage is skill vs. luck. If it’s mostly luck, then the commentariat is wasting their time, which suggests I should unfollow. But if it’s mostly skill, then maybe I should take heed and start supporting their Patreons.

As is my wont, this meant that I had to do a reliability analysis of college completion percentage (Comp%). And then I figured, I might as well assess the reliability of college Yards per Attempt (YPA) and Adjusted Yards per Attempt (AY/A) too. ((For the unaware, I couldn’t assess ANY/A because the NCAA — in their infite wisdom — still counts sacks as running plays to this day.)) What follows is a post detailing these statistical inquiries, their results, and their implications, both theoretical and applicable to the 2018 draft.

Methods

It’s been a while, and also some of you may be new to this, so here are the details of how I went about performing this reliability analysis:

  1. I collected data for all Football Bowl Subdivision (FBS) QBs that had at least 8 games with 14 or more pass attempts from 2000 to 2017. ((Data courtesy of and qualifying threshold per Sports Reference.))
  2. To control for team effects, I included only those QBs that played 8+ games for the same team.
  3. Starting with QBs that played 8+ games, I randomly selected two sets of 4 games for each QB, and calculated their Comp%, YPA, and AY/A in both sets.
  4. For both of these metrics, I calculated its split-half correlation (r) between the two randomly-selected sets of games.
  5. I performed 25 iterations of Step 4 so that r converged.
  6. I repeated Steps 3-5, increasing the QB inclusion criteria in 8-game intervals, from 8+ games all the way to 48+ games. ((Long-time readers will notice that this step proceeds through 72+ games in an NFL context. The reason for stopping at 48+ games is that it’s impossible to play 72 games in college , and 48 games approximates a four-year starter.))
  7. For each “games played” group, I calculated the number of games at which the variance explained in each metric, R2, would mathematically equal 0.5. ((The formula is (Observations/2)*[(1-r)/r].))
  8. I calculated the True Comp%, True YPA, and True AY/A for a hypothetical QB that’s had an observed performance of 60.0% Comp%, 9.75 YPA, and 10.50 AY/A through 8, 16, 24, etc. number of games. ((The formula is [(Observed Performance * Observations) + (League-Average Performance * Stabilization Point)] / (Observations + Stabilization Point) ))
  9. I calculated a weighted average of the results from Steps 7 and 8. ((Weighted by group size.))

Results

First up, I’ll answer the question, “How many games-worth of stats do we need to see before an FBS QB’s YPA represents 50 percent signal (aka skill) and 50 percent noise (aka luck)?”

Game SplitnrR2 = 0.50Avg Y/AObs 9.75 Y/A
41,0860.29107.178.46
88840.45107.218.48
125670.54107.368.56
163490.58127.438.59
201570.62127.438.59
24120.65137.758.75
Wtd Average107.268.51

For those unfamiliar or who have (understandably) forgotten, here’s how to read the above results table. The “bottom line” statistic is the stabilization point, which is displayed in the “R2 = 0.50″ column of the “Wtd Average” row. In this case, said statistic reveals that college YPA takes 10 games to stabilize, which means a YPA based on 10 games represents 50 percent signal and 50 percent noise. Alternatively, it means that said YPA should be regressed exactly halfway towards the mean to determine the FBS QB’s “true” (i.e., noise-independent) YPA.

The rest of the table’s “Wtd Average” row shows that a hypothetical FBS QB with a YPA of 9.75 after 10 games has a True YPA of 8.51, which is exactly halfway between 9.75 and 7.26 (i.e., the overall average YPA in my sample).

And if we want, we can increase granularity by exploiting the knowledge that the QBs in my sample averaged 25.6 attempts per game. ((720,834 attempts across 28,174 QB games)). Doing so reveals that YPA takes 264 attempts to stabilize. Interestingly, this is lower/quicker than the 396-attempt stabilization point I found for NFL YPA.

So what happens when we adjust YPA for touchdowns and interceptions so as to produce AY/A?

Game SplitnrR2 = 0.50Avg AY/AObs 10.50 AY/A
41,0860.28106.788.64
88840.43116.878.69
125670.52117.138.81
163490.54147.248.87
201570.59147.268.88
24120.7687.839.17
Wtd Average116.958.73

Given how connected AY/A is to YPA mathematically, the similarity of results makes perfect sense. This extends to attempts too, as the mathematical translation to 284 attempts is negligibly different from YPA’s 264 attempts.

Furthermore, if adjusting for touchdowns and interceptions doesn’t have a meaningful impact on reliability with respect to FBS QBs, then why even have a stat that adjusts FBS QB YPA for touchdowns and interceptions? If I had to guess, this situation is the (understandable) result of early FBS analytics taking its cue from NFL analytics just as early NFL analytics took its cue from MLB analytics (e.g., Pythagorean wins).

Moving on, here are my reliability analysis results for Comp%:

Game SplitnrR2 = 0.50Avg Comp PctObs 60.0% Comp Pct
41,0860.38658.6%59.3%
88840.56658.9%59.5%
125670.64759.9%60.0%
163490.71660.3%60.2%
201570.75760.4%60.2%
24120.88362.5%61.2%
Wtd Average759.2%59.6%

Lo and behold, FBS Comp% stabilizes faster than FBS YPA and FBS AY/A. In terms of games, the table shows 7 versus the previous 10 and 11. And translating this result into pass attempts, 167 is considerably lower than 264 and 284. Therefore, we can conclude that Comp% is the “stickiest” of these three stats; the most reflective of statistical signal vis-a-vis noise.

However, two caveats apply. First, as Josh Hermsmeyer has shown, and Bill Barnwell has applied to this year’s crop of FBS QBs, Comp% must be placed in the context of target depth (i.e., adjusted for air yards). My analysis did not do so, which is a clear limitation. That said, as you’ll see shortly, insofar as using this result to attach a “True Comp%” number to each QB in the upcoming draft, the rankings are similar enough to Bill’s to render this caveat moot (for now).

Second, and I want to make this point abundantly clear, I’m reporting a reliability analysis here, not a validity analysis. Just because FBS Comp% is “sticky,” doesn’t necessarily mean it predicts success at the NFL level. It’s said that “reliability sets the ceiling for validity.” Applied to our current statistical consideration(s), this means that FBS Comp% can’t predict NFL Comp% better than FBS Comp% predicts itself from one game to the next. Therefore, the best I can say based on my results is that FBS Comp% should be — must be — converted to True Comp% in predictive models. ((See Footnote 3 below.))

Speaking of which, my last topic of discussion wishes to take a 30,000-foot view of QB projection. In the results above, it’s clear that FBS Comp% is the most reliable metric of the three. Astute observers will have also noticed that it increases commensurate with the number of games played (i.e., as the “Game Split” column increases). Hell, it does for YPA and AY/A too, albeit to a lesser extent. Of course, this makes sense, as we would hope QBs get better on average with experience.

All of which is to say that the original idea underlying Football Outsiders’ original QB projection system — the Lewin Career Forecast — seems sound from a reliability perspective, even if it hasn’t stood the test of time from a predictive perspective. The two most important variables identified by Lewin over a decade ago were Comp% and Games Started. Perhaps the cutoffs — 60% Comp% and 37 games started — that arose later were arbitrary. Perhaps the 65% R2 he reported in 2007 suffered from overfitting. Nevertheless, those devils in the details should not obscure the overarching angel of a fact that the combination of Comp% and sample size (Read: Experience) reliably indicates FBS QB skill.

Are they predictive? This analysis can’t answer that. Reliability sets the ceiling on validity, so the point here is that these variables were a decade ago — and continue to be — a good place to start.

Statistics wonkery out of the way, below are True Comp%, True YPA, and True AY/A for the 16 FBS QBs who currently have a draft round designation by NFLDraftScout.com:

PlayerNFLDS RdTrue Comp PctRkTrue YPARkTrue AY/ARk
Josh Allen156.9%167.65127.4812
Baker Mayfield169.8%110.62111.861
Sam Darnold164.9%68.5468.727
Josh Rosen160.9%107.98107.9911
Lamar Jackson157.0%158.3388.489
Mason Rudolph263.2%89.4129.872
Mike White3-466.4%38.7549.354
Luke Falk468.3%27.05137.3513
Chase Litton5-660.8%116.96146.9914
Nic Shimonek666.4%48.0498.558
Kurt Benkert6-757.5%146.29166.3216
Logan Woodside6-765.1%59.0239.653
Riley Ferguson7-FA63.1%98.6859.275
J.T. Barrett7-FA63.5%77.79118.3910
John Wolford7-FA59.7%136.90156.3815
Danny Etling7-FA59.7%128.4378.846

Again, this post was on reliability (i.e., measurement), not validity (i.e., prediction), but I’m compelled to opine based solely on the former:

  • Baker Mayfield laps the field in all three “true” stats, as he did based on Barnwell’s aforementioned air yards-adjusted stats.
  • Josh Allen’s “true” stats aren’t commensurate with a No. 1 pick.
  • Lamar Jackson’s and Josh Rosen’s “true” stats resemble Allen’s more than Mayfield’s.
  • Mason Rudolph.

DT : IR :: TL : DR

It takes about half-a-season for completion percentage to stabilize for college QBs, whereas it takes about a full season for YPA and AY/A. Also, the more games a college QB plays, the “stickier” these stats get. This finding doesn’t necessarily mean college completion percentage is predictive of NFL success, especially when it’s not adjusted for depth of target (aka air yards). That’s because predicting NFL success is a validity question, not a reliability question. Based on my results, the answer to said question is to use True Comp%, True YPA, and True ANY/A in predictive models rather than their raw counterparts.

Bookmark the permalink.

Leave a Reply

Your email address will not be published. Required fields are marked *