
Forecast how long an LLM reply will run, before Enter
Token Forecaster puts a range on an LLM reply before you press Enter: the usual length, and a worst case that held for 90.6% of 4,146 unseen calls. While the reply streams, it shows whether it is running long. It only watches and never changes the request. Learns from your local history. Open source, MIT.
Token Forecaster estimates the length of an LLM reply before submission, providing a typical and worst-case duration based on historical data from 4,146 calls. It monitors the reply duration in real-time without altering the request and is available as an open-source tool under the MIT license.
Scored deterministically. Only candidates that fire a story trigger are sent to a model, so this one has no written angle.
Gaps in our data, not findings about the product. Their weight is redistributed across the 5 we did measure.
A source that found nothing is a measurement. A source that has not run is a gap. Neither means the launch lacks the thing.