AI FTP versus other operationalizations of FTP

There seems to be a sentiment on this forum that the new TrainerRoad AI FTP is less related to the underlying theoretical construct of Functional Threshold Power than other operationalizations. Is there any data (beyond anecdote) to suggest this is empirically true?

Imagine we have 1000 people do a 20-minute FTP test, an 8-minute FTP test, and a Ramp test. And we also compute their TrainerRoad AI FTP. Then we calculate the percentage differences between each pair of tests (20 - 8, 20 - Ramp, 20 - AI, 8 - Ramp, 8 - AI, and Ramp - AI). Finally, we plot a histogram of the differences for each pair of tests. What would these histograms look like? Perhaps they would be approximately Gaussian? What would the means and standard deviations be? Would the means for comparisons with AI FTP be any further from zero? Would the standard deviations for comparisons with AI FTP be any larger?

If the histograms involving the AI looked about the same as those between tests, would folks accept that while AI FTP is different than the other measures it is no better or worse as an approximation of the underlying construct?

There are quite a few major obstacles that, in the eyes of many who care, would immediately render such an exercise moot. A few that come to mind are:

  • What is the definition of FTP? If it is e. g. power at MLSS, then I don’t think TR has the data to determine that.
  • You absolutely need to figure out the uncertainties in everything. I’ve seen claims that were attributed to Coggan that FTP test results may vary by 3–5 % depending on an athlete’s physical and mental state on the day of testing. You don’t have to agree to those specific numbers, but we all know that this is likely an effect you need to control for. And at the end, it may turn out to be a wash, i. e. a lot of FTP test protocols may give you the same number within the margin of error.
  • Many exercise scientists regard Critical Power as essentially identical to FTP for that reason, too.
  • If you instead pick best power at a specific duration, then you need to make sure you have good data, i. e. you really have all-out efforts for that duration on file.
  • The distribution need not be Gaußian and have longer tails with slower decay or are asymmetric around the modal point.
  • At the end, what would be the value? Coggan is against the 20-minute test, which was the benchmark amongst practitioners for many years until the ramp test came about. (I’m not stating my opinion on either, I have never done a 20-minute test in my life, just saying that if you are an athlete or coach that “grew up” in the early days of power meters, you’d incorporate 20-minute tests in your training.)
  • AI FTP could, at the end of the day, still result in better training even though it is off from “the true” FTP.

I appreciate all of your great points.

This makes for a great theoretical discussion, but as your other points make clear, whatever the definition is, that isn’t what any of the typical tests or the AI FTP are directly measuring.

Yes! This is sort of the key point I was trying to get at. People seem to be upset about how TR is going about modeling FTP, but without anything more than anecdotal evidence that their approach is doing any worse than any of the typical tests. They all presumably have their strengths and weaknesses and work better for some people more than others.

I was trying to provide a thought experiment about the sort of data that might convince people that AI FTP is no better or worse than the typical FTP tests for approximating the underlying construct of FTP.

I do think that people have come to rely on FTP as a portable measure that they can carry with them from one context to another to guide their activities. For people who are only using it within TR, I agree that it doesn’t really matter. But it seems like a lot of people use it across multiple apps, with IRL coaches, out on the road, etc… So, for those folks, knowing that the AI FTP is estimating something aligned with what the typical tests are measuring would be very reassuring.

If you want to make any comparison, you have to choose a standard that you are comparing others to. Otherwise, you will not even be able to ask the question you seem to be interested in.

But that’s a different question: estimating the variability of one single FTP test protocol is definitely interesting, but ultimately, a different question.

The conclusion, though, may be highly relevant for a comparison, because the error bars it establishes are crucial when one wants to decide whether you can distinguish two FTP tests from one another.

Moreover, there are other sources of variability, e. g. when you use the trainer’s power meter for workouts and another power meter for outside rides.

Hmmm, I think you can ask (and TR has asked and answered) this different question: which method to determine the “FTP value” used to scale workouts works best within TR.

TR’s data is apparently very clear: TR’s new AI FTP works best within TR. But that’s again a very different question.