AI FTP and the cone of uncertainty

Thanks for engaging.

I don’t believe any of us have the full details of what AI technologies are being employed by TR, if we are through the Beta programme then I’m sure we’ve been asked not to share publicly.

The concern you are describing can be summarised by saying “how can we believe the output from TR to be trustworthy?”. You have analogised this to hurricane uncertainty as below:

A reasonable analogy is hurricane forecasting. You are the hurricane. Your AI FTP is the position of the storm. Your upcoming workouts are the terrain the hurricane will pass over

Whilst you are describing uncertainty, I don’t think you are correct in your analogy because with hurricanes it is not that we don’t know how hurricanes behave, it’s the chaos in the system that makes it hard to predict.

Per se - This is not equivalent to training data or the TR ML analysis and prediction.

The outcome you are looking for - include a ‘confidence level’ in the prediction- is reasonable, but the rationalisation and path to get there is not.

Kind of difficult to do over forum posts, but I hope that makes sense. :slightly_smiling_face: I’ve left out non-determinism because that’s a bigger issue.

I don’t think you are an outlier. I would bet that for most people this is working really well.

Good luck with the race!

I think the analogy still holds up because in each case there’s undefined uncertainty that contributes to the range of possible outcomes. The outcome for the hurricane is its position in the future. For TR, it’s the wattage of your FTP in the future.

Hurricanes have a lot of chaos contributing to the storm path, which is what makes their future location a range of possibilities that gets wider the further in the future you look. The further out, the more factors there are that have time to change the path of the hurricane

TR workouts have less in the way of chaotic factors, but it’s not zero. Each workout has the opportunity to go well or poorly based on one’s health, recovery, overall fitness, sleep, nutrition, mental state and whether or not the podcast you’re listening to is any good. And for that workout you can finish as indicated, or crank it up, or crank it down, or fail, or pause for breath, skip it entirely, change out for a different workout, etc.

I think the range of factors that come into play with any given workout are a decent analog of the factors that contribute to the storm path of a hurricane. And like a hurricane, early changes change the range of possibilities moreso than late changes. A storm isn’t as likely to make a significant course correction over 24 hours, but over 7 days it can change radically. Likewise, your AI FTP prediction isn’t likely to vary much with only 1 workout left before FTP detection, but might change significantly in the first week of a block if your workouts aren’t as “expected” in either positive or negative ways.

Feels a little unhelpful to say, “when you know, you know,” but that’s how it feels a bit to me. I think my barometer is that I don’t want to be shattered by a workout so I’m looking for something that rides the line between my response of Hard and Very Hard. If you get through the final rep of the final set and think, I could do another rep, but I don’t really want to or feel the need to, then I think it’s enough. If it’s too easy, it’s not the right workout, but also back to the original point, TR can adjust that difficulty without changing your FTP.

That might be a way to get the uncertainty without running a whole bunch of simulations. Can they identify the couple of factors that drive the variability and then cluster us users into groups based on those variables.

Maybe @Jonathan and @Nate_Pearson can have a follow up podcast where they walk through the variables that are driving variability in the results.

I don’t think they are assuming you respond optimally, they are assuming you respond with the most likely response. I think that is slightly different. Although they have made some assumptions that people respond rationally and consistent with themselves.

Nate has said repeatedly that their model is deterministic.

Cool. Would be great to see a quote or link on that?

I think the thing not clicking for people is the fact that the way the hurricane and TR models work is by taking data from right now and run the model once to predict the days one small time increment in the future. Then, take those predicted results as the new input to the model and run it again to get the next small time increment into the future. Keep doing that until you get to a final result. This is why every time you tweak your planned or actual training you have to wait for TR to calculate each future workout one by one. They can’t predict the 28th day until they have or predict data for days 1-27.

This means even small errors, for example <1% per simulation, add up a lot when you use those errors over and over again (28 or more simulations to get to 28 days into the future).

That’s why it’s ideal to run the simulations many times with assuming your estimate was slightly high or low at each step to see how big those little errors add up. Then you might use the most occurring results as your highest confidence prediction. This is a simplified explanation of Monte Carlo technique that’s often used.

However, it’s expensive and time consuming to perform.

It was in one of the recent podcasts with Nate.

It’s deterministic as in if you feed it the exact same data, you’ll get the exact same result. However, that doesn’t mean that anyone at TR can explain why the model gave that result. That’s the problem with neural networks. You don’t know everything happening inside the black box. You don’t know the reasons for the result, just the result.

We are saying the same thing. Optimal == most likely response. But if the most likely is only 60%, then do the math and assuming a 60% chance happens 10, 20, more times and the outcome will almost never happen. Which is my point.

Interesting comments.

I rekon it’s a bit like steering an oil tanker. You just need a small nudge in either direction early on to keep the correct trajectory. If you drift off course (missed /failed workouts, for whatever reason) and don’t immediately correct, then it becomes harder and harder to get back on track, until you reach the point of no return!

So my approach is…I open up my iPad, fire up the App, check the workout power profile vs my recent best efforts, so I know what to expect, and then crack on.

Not once has the prescribed workout felt harder / easier than the AI prediction, and I haven’t failed a single workout…yet!

I ‘feel’ stronger and my FTP (FWIW) is edging up. currently 278 and predicting 280 in 3 weeks time. For context, I’m 60 years old, 75kg, and following a ‘balanced’ Masters Plan, but with 2 x 1 hour strength sessions / week on non-cycling days. I do all the interval stuff on the turbo and the endurance rides outdoors. Seems to work for me.

Very happy bunny! :smiling_face_with_sunglasses:

Right, but was your normalized or average power higher in year 3 for this event? Or was the improved time due to one of a dozen other factors that TR isn’t selling: wind, temp, tire selection, pressure, total weight bike+rider, road surface and so on.

This discussion is great and “gives us furiously to think”.

But.

For the vast majority of TrainerRoad users, a predicted range, or an accuracy tolerance, would simply over-complicate and confuse.

Of course I’d love to know more. I’m Mr Statto. I live for spreadsheets. Do I need to know more?

Hmmm.

Lots of posts about earned predicted gains being taken away. So I think either a more conservative initial estimate or the range would be helpful.

The sensitivity is nice if you are using as a single if something will help or hurt.

My $0.02,

What’s missing is an after action report that provides insight into why the Detected AI FTP did or did not meet the initial Predicted AI FTP.

Something like: Your detected AI FTP was 5 watts below your predicted AI FTP because

  • Your compliance to the plan was only 75%
  • You manually increased the load, which appears to have added fatigue leading you to under-perform on key workouts
  • etc.

Now something like would help athletes understand what they are doing that is negatively (which is what it seems like from the forum posts) impacting their detected compared to their predicted AI FTP.

Would be helpful. I’d even be interested in:

“You did an extra 4h group ride because the weather was perfect and you like riding outside with your friends more than the 90min @65% that was prescribed for the day. Even though you completed the next 2 weeks of interval workouts perfectly, never rating them harder that predicted, the algorithm states that the extra TSS of that group ride 2 weeks ago means you lost aerobic fitness at the end of the block. Minus 8W off AI FTP.”

More prosaic than my example, but the same exact idea. Something like this would really help people be intentional about trade-offs: in your example, how much do I value riding outside on a gorgeous day with friends versus doing an indoor ride to get a higher predicted FTP? For me, that’s an easy trade-off: the outside group ride with friends. But I don’t race.

NO! The forum will be filled with “WHY IS MY RANGE GOING DOWN"!?

No matter what TR does they’re screwed if they do and screwed if they don’t.

Probably. I guess the real solution is education by TR on what moves the prediction.