Quick Answer
timesfm-3.0-pytorch (google) is the top pick as of September 2026 because it is the newest release among the candidates, with no public benchmark yet to establish comparative demand-forecasting accuracy [3][5][4][10][7][9][1]. In order, the ranking is timesfm-3.0-pytorch (google), granite-timeseries-patchtst-fm-r2 (ibm-granite), granite-timeseries-ttm-r3 (ibm-granite), chronos-2-small (autogluon), chronos-2 (amazon), chronos-2 (autogluon), and timesfm-2.5-200m-pytorch (google) [3][5].
Key Takeaways
- Google’s
timesfm-3.0-pytorch (google)ranks first because its publication date, August 24, 2026, is the newest among the eligible releases; its license istimesfm-non-commercial-license-v1.0.[3] - IBM Granite’s
granite-timeseries-patchtst-fm-r2 (ibm-granite)ranks second, with 385M parameters and theopenmdw-1.0license.[5] IBM Granite’sgranite-timeseries-ttm-r3 (ibm-granite)ranks third, with 1M parameters and the Apache-2.0 license.[4] - AutoGluon’s
chronos-2-small (autogluon)ranks fourth, with F32 weights of 0.1 GB.[10] Amazon’schronos-2 (amazon)ranks fifth, and AutoGluon’schronos-2 (autogluon)ranks sixth; each has F32 weights of 0.5 GB.[7][9] All carry the Apache-2.0 license.[10][7][9] - Google’s
timesfm-2.5-200m-pytorch (google)ranks seventh as an established pick: membership among its family’s three most-downloaded models lets it bypass the twelve-month recency gate, without qualifying it for first place through that exemption.[1] - Local deployment comparisons include F32 weight sizes of 1.3 GB for Google’s ranked-first checkpoint and 1.5 GB for IBM Granite’s ranked-second checkpoint.[3][5] Treat those figures as weight sizes, not verified GPU or system RAM requirements.
- Ranking note: no public benchmark yet scores these candidates against each other, so the ordering is newest release first, then downloads.[3][5][4][10][7][9][1] Scope admits only time series forecasting models tagged
time-series-forecastingon Hugging Face from labs with a published paper or leaderboard record, or other publishers above 10,000 downloads in the last 30 days.[2][6][8][11][4]
How do local demand forecasting models compare on specifications and benchmarks?
| Model | Org | Params | Quant/VRAM | Released (date) | License | Key benchmark (date) |
|---|---|---|---|---|---|---|
| google/timesfm-3.0-pytorch [3] | google [3] | 331M [3] | F32 weights: 1.3 GB [3]; VRAM: not published | 2026-08-24 [3] | timesfm-non-commercial-license-v1.0 [3] | no public benchmark yet |
| ibm-granite/granite-timeseries-patchtst-fm-r2 [5] | ibm-granite [5] | 385M [5] | F32 weights: 1.5 GB [5]; VRAM: not published | 2026-08-07 [5] | openmdw-1.0 [5] | no public benchmark yet |
| ibm-granite/granite-timeseries-ttm-r3 [4] | ibm-granite [4] | 1M [4] | F32 weights: 0.0 GB (reported rounded value) [4]; VRAM: not published | 2026-05-21 [4] | Apache-2.0 [4] | no public benchmark yet |
| autogluon/chronos-2-small [10] | autogluon [10] | 28M [10] | F32 weights: 0.1 GB [10]; VRAM: not published | 2025-12-03 [10] | Apache-2.0 [10] | no public benchmark yet |
| amazon/chronos-2 [7] | amazon [7] | 119M [7] | F32 weights: 0.5 GB [7]; VRAM: not published | 2025-10-30 [7] | Apache-2.0 [7] | no public benchmark yet |
| autogluon/chronos-2 [9] | autogluon [9] | 119M [9] | F32 weights: 0.5 GB [9]; VRAM: not published | 2025-10-06 [9] | Apache-2.0 [9] | no public benchmark yet |
| google/timesfm-2.5-200m-pytorch — established pick [1] | google [1] | 231M [1] | F32 weights: 0.9 GB [1]; VRAM: not published | 2025-09-02 [1] | Apache-2.0 [1] | no public benchmark yet |
Which forecasting models should you evaluate for local demand forecasting?
1. timesfm-3.0-pytorch
timesfm-3.0-pytorch by google ranks first on release recency, with a Hugging Face publication date of August 24, 2026.[3] The ordering is newest release first, then downloads; its position does not establish superior demand-forecasting accuracy. Benchmark status: no public benchmark yet, so there is no dated public task score to report.
The checkpoint contains 331 million parameters and lists F32 weights of 1.3 GB.[3] For local hardware planning, treat that weight size as a starting point, not a total memory requirement. A specific RAM capacity or GPU fit would require an estimate; the listed specifications do not establish either. Measure memory use with your intended forecast workload before selecting hardware.
Use the checkpoint as a candidate for local demand-forecasting evaluation, checking forecast quality on your own historical demand before deployment. The licensing caveat is its listed timesfm-non-commercial-license-v1.0 license.[3] Review those terms against your intended use before adoption.
2. granite-timeseries-patchtst-fm-r2
granite-timeseries-patchtst-fm-r2 from ibm-granite ranks second under the ordering rule “newest release first, then downloads”: its Hugging Face publication date is August 7, 2026, behind Google’s timesfm-3.0-pytorch (google), published August 24, 2026.[5][3] Its benchmark status is “no public benchmark yet,” so the position does not establish superior demand-forecasting accuracy.
The checkpoint has 385M parameters, with F32 weights listed at 1.5 GB, and carries the openmdw-1.0 license.[5] The weight size is a storage reference, not a total runtime-memory requirement. Hardware requirements for local inference are unspecified; measure memory consumption with your intended workload before choosing a GPU or committing RAM.
Use the model as a candidate for a local demand-forecasting evaluation against your existing baseline. Test on your own demand history before selecting it for deployment. The practical caveat is the absence of a public benchmark result: release recency establishes its position here, but does not demonstrate accuracy on your inventory or sales data.
3. granite-timeseries-ttm-r3
granite-timeseries-ttm-r3 from ibm-granite ranks third under the ordering rule: newest release first, then downloads.[3][5][4][10] Published on 2026-05-21, it follows the Google and ibm-granite releases ahead of it and precedes the next-ranked AutoGluon release.[3][5][4][10] Its position reflects release timing; no public benchmark yet establishes its demand-forecasting accuracy against the other candidates.
The checkpoint has 1M parameters and lists Apache-2.0 as its license.[4] For local deployment, use that parameter count as a starting point for evaluation. A specific CPU, GPU, or RAM requirement cannot be established from the listed specifications. Confirm runtime support and measure peak memory on your intended hardware before planning deployment.
Consider it for a demand-forecasting pilot where model footprint matters. Evaluate forecasts against your existing baseline using your own demand history and forecast horizon. The caveat is accuracy: its ranking supplies no demonstrated advantage for your workload.
4. chronos-2-small
chronos-2-small from autogluon ranks fourth under the ordering “newest release first, then downloads,” with a Hugging Face publication date of December 3, 2025.[10] Its benchmark status is “no public benchmark yet,” so its position does not establish demand-forecasting accuracy relative to the other candidates.
The checkpoint has 28 million parameters and 0.1 GB of F32 weights, with Apache-2.0 listed as its license.[10] For local deployment, treat that weight footprint as a starting point for memory planning, not a complete hardware requirement. A specific CPU, GPU, minimum RAM or minimum VRAM requirement is not established; measure total memory use with your intended workload.
Consider it for a local demand-forecasting pilot where keeping the weight footprint small matters. Validate forecasts against your own demand history before choosing it for operational use. The caveat is accuracy uncertainty: release recency and parameter count do not demonstrate suitability for your demand patterns.
5. chronos-2
chronos-2 by amazon ranks here under “newest release first, then downloads,” with a first publication date of 2025-10-30.[7] Its release falls between autogluon’s chronos-2-small, published on 2025-12-03,[10] and autogluon’s chronos-2, published on 2025-10-06.[9] The position reflects release timing, not demonstrated demand-forecasting accuracy: no public benchmark yet.
The checkpoint has 119M parameters and 0.5 GB of F32 weights, with an Apache-2.0 license.[7] For local hardware planning, use the weight size as a starting point, then measure runtime memory with your intended batch size and forecast horizon before choosing a machine. Do not treat the weight footprint as a total RAM or GPU memory requirement.
Use the checkpoint as a candidate for local demand-forecasting backtests on your sales history. Compare its forecasts against your existing baseline across replenishment horizons and seasonal periods before adopting it; the ranking alone does not establish suitability for your inventory decisions.
6. chronos-2
chronos-2 by autogluon ranks sixth under the ordering rule “newest release first, then downloads”: its Hugging Face publication date is 2025-10-06, earlier than amazon’s chronos-2 release on 2025-10-30.[9][7] The checkpoint has no public benchmark yet, so its position does not establish demand-forecasting accuracy. Its 8,232,206 downloads over the preceding thirty days indicate adoption, not forecast quality.[9]
The model contains 119M parameters, with published F32 weights occupying 0.5 GB, and lists Apache-2.0 as its license.[9] For local hardware planning, treat that weight size as a starting point rather than a total memory requirement. A specific GPU or RAM recommendation would require a runtime measurement or an explicitly labelled estimate.
Use this checkpoint as a local evaluation candidate for your demand histories. Compare forecasts against your existing planning baseline across replenishment horizons and seasonal changes. The practical caveat is the absence of a public benchmark establishing its performance on demand forecasting; download volume cannot resolve that uncertainty.
7. timesfm-2.5-200m-pytorch
timesfm-2.5-200m-pytorch from google ranks seventh as an established pick.[1] Its first Hugging Face publication was September 2, 2025, and it qualifies through the exemption for the family’s three most-downloaded models.[1] The ordering is newest release first, then downloads; its established status does not override its publication date. The newer google timesfm-3.0-pytorch was published on August 24, 2026.[3]
The checkpoint has 231M parameters, a 0.9 GB F32 weight file, and an Apache-2.0 license.[1] For local hardware planning, use that weight-file size as a starting point, then measure runtime memory before choosing a GPU or allocating system RAM. A specific GPU or RAM requirement is not established here.
Use it as an established baseline when evaluating demand forecasts locally, particularly when comparing checkpoints by license and deployment requirements. The caveat is accuracy evidence: no public benchmark yet establishes its demand-forecasting performance against the ranked candidates, so validate it on your own demand history before adoption.
How are forecasting models ranked when no public benchmark scores the candidates?
When no public benchmark scores the candidates, the ranking uses newest release first, then downloads; Google’s TimesFM 3.0 (timesfm-3.0-pytorch (google)) precedes IBM Granite’s granite-timeseries-patchtst-fm-r2 (ibm-granite) on publication date.[3][5] Every unscored candidate is labeled “no public benchmark yet.”
The ranking admits only time series forecasting models from labs with a published paper or leaderboard record, or other publishers above 10,000 downloads in the last 30 days, using Hugging Face’s time-series-forecasting pipeline tag.[1][3][4][5][7][9][10] Downloads serve only as a tiebreak; popularity does not establish demand-forecasting accuracy.
The recency gate normally requires a release within the last 12 months.[3][4][5][7][9][10] Google’s TimesFM 2.5 (timesfm-2.5-200m-pytorch (google)) is an established pick: its September 2, 2025 publication falls outside that window, but its position among the family’s three most-downloaded models grants an exemption.[1] Established picks follow the same ordering rule and cannot take first place through that exemption.[1]
Google’s TimesFM 3.0 takes first place because its August 24, 2026 publication is the newest eligible release, not because of a demonstrated demand-forecasting benchmark advantage.[3][5]
A leaderboard may compare only candidates it actually scores, and an older scored model cannot outrank a newer unscored model. Model-card benchmark results must be labeled “self-reported.”
Parameter counts and weight sizes inform deployment planning, not ranking. Hardware-fit claims require explicit estimates. License identifiers are reported as listed; a license name alone does not establish deployment permissions.
What do published weight sizes tell you about local hardware requirements?
Published weight sizes provide a starting point for local storage planning; use them as weight footprints, not as complete RAM or GPU memory requirements.
Google’s timesfm-3.0-pytorch (google) lists F32 weights of 1.3 GB.[3] IBM Granite’s granite-timeseries-patchtst-fm-r2 (ibm-granite) lists F32 weights of 1.5 GB.[5] Use those figures when budgeting storage for the checkpoints. Neither figure establishes a specific GPU recommendation or total system memory requirement.
IBM Granite’s granite-timeseries-ttm-r3 (ibm-granite) lists 1M parameters and an F32 weight size of 0.0 GB.[4] Do not interpret that displayed size as a zero-memory requirement. The entry does not provide a sufficiently precise weight footprint for hardware sizing.
AutoGluon’s chronos-2-small (autogluon) lists F32 weights of 0.1 GB.[10] Amazon’s chronos-2 (amazon) and AutoGluon’s chronos-2 (autogluon) each list F32 weights of 0.5 GB.[7][9] Google’s timesfm-2.5-200m-pytorch (google) lists F32 weights of 0.9 GB.[1] Those are checkpoint weight figures; a claim that any checkpoint fits a particular machine would require a separate estimate or measurement.
For demand forecasting, validate the intended deployment with your forecast horizon, batch size and actual workload before purchasing hardware. Measure peak memory and inference time during that evaluation. Keep published weight size separate from measured runtime requirements, and treat a named CPU, GPU or RAM configuration as unverified until you have supporting deployment evidence.
Which licenses are listed for the forecasting checkpoints?
The forecasting checkpoints list Apache-2.0,[1][4][7][9][10] openmdw-1.0,[5] and timesfm-non-commercial-license-v1.0.[3] The listed license belongs to the individual checkpoint, so keep the repository identifier attached to each license entry during evaluation.
Google’s timesfm-3.0-pytorch (google) lists timesfm-non-commercial-license-v1.0.[3] Google’s timesfm-2.5-200m-pytorch (google) lists Apache-2.0.[1] Record those entries separately when comparing the Google checkpoints; do not carry a license label forward from a previously evaluated release.
IBM Granite’s granite-timeseries-patchtst-fm-r2 (ibm-granite) lists openmdw-1.0.[5] IBM Granite’s granite-timeseries-ttm-r3 (ibm-granite) lists Apache-2.0.[4] Keep the exact license identifier in your deployment inventory rather than replacing it with a general description such as “open weights.”
Amazon’s chronos-2 (amazon) lists Apache-2.0.[7] AutoGluon’s chronos-2 (autogluon) and chronos-2-small (autogluon) also list Apache-2.0.[9][10] Preserve the publisher and full repository name when documenting which checkpoint your demand forecasting application uses.
For deployment review, treat these entries as license identifiers. Read the applicable license text before deciding whether your intended use, modification or redistribution is permitted. Record that decision alongside the selected checkpoint and retain any applicable notices. Avoid turning a shared license label into an assumption that every checkpoint has already passed your organization’s approval process.
Frequently Asked Questions
Which forecasting model should I evaluate first for local demand forecasting?
Google’s timesfm-3.0-pytorch (google) ranks first because its publication date makes it the newest release among the eligible candidates.[3][5][4][10][7][9][1] Published on August 24, 2026, the checkpoint has 331M parameters and lists timesfm-non-commercial-license-v1.0.[3] Its benchmark status is “no public benchmark yet”; the ranking therefore does not establish superior demand-forecasting accuracy.[3] Evaluate it against your existing forecasting baseline before selecting it.
How are the local forecasting models ranked?
Ranking note: newest release first, then downloads. Order: Google’s timesfm-3.0-pytorch (google),[3] IBM Granite’s granite-timeseries-patchtst-fm-r2 (ibm-granite),[5] IBM Granite’s granite-timeseries-ttm-r3 (ibm-granite),[4] AutoGluon’s chronos-2-small (autogluon),[10] Amazon’s chronos-2 (amazon),[7] AutoGluon’s chronos-2 (autogluon),[9] Google’s timesfm-2.5-200m-pytorch (google)—an established pick.[1] Scope admits only models tagged time-series-forecasting from labs with a published paper or leaderboard record, or other publishers above 10,000 downloads in the last 30 days.[1][3][4][5][7][9][10] Downloads serve only as a tiebreak; parameter count does not determine order.
Do public benchmarks show which model forecasts demand more accurately?
Every ranked checkpoint has the status “no public benchmark yet” for this comparison; no public benchmark scores these candidates against each other.[1][3][4][5][7][9][10] Publication dates determine the ordering, rather than demonstrated demand-forecasting accuracy. For your evaluation, compare forecasts on held-out demand history using the same forecast horizon, available inputs and business error measure. Treat any model-card benchmark you subsequently consult as self-reported.
How much memory do I need to run these models locally?
Published weight sizes provide a starting point, but do not establish a verified RAM or GPU requirement. AutoGluon’s chronos-2-small (autogluon) lists F32 weights of 0.1 GB,[10] while Google’s timesfm-3.0-pytorch (google) lists 1.3 GB.[3] IBM Granite’s granite-timeseries-patchtst-fm-r2 (ibm-granite) lists 1.5 GB.[5] Measure peak memory with your intended workload before choosing hardware; avoid treating weight-file size as a complete deployment requirement.
Which licenses do these forecasting checkpoints list?
Google’s timesfm-3.0-pytorch (google) lists timesfm-non-commercial-license-v1.0,[3] and IBM Granite’s granite-timeseries-patchtst-fm-r2 (ibm-granite) lists openmdw-1.0.[5] The remaining ranked checkpoints list Apache-2.0.[1][4][7][9][10] License names alone do not establish the permissions applicable to your deployment. Read the exact checkpoint’s license before approving its use, including for internal demand planning, redistribution or a customer-facing forecasting service.
Why include an established checkpoint alongside newer releases?
Google’s timesfm-2.5-200m-pytorch (google) is an established pick: it qualifies through the family’s download-based exemption from the 12-month recency gate.[1] Published on September 2, 2025, it has 231M parameters and lists Apache-2.0.[1] Established picks follow the same ordering rule and cannot take first place through the exemption. Include this checkpoint as an evaluation baseline without assuming its popularity demonstrates forecasting accuracy.
Sources
- google/timesfm-2.5-200m-pytorch model card (Hugging Face) — 2026-09-25
- A decoder-only foundation model for time-series forecasting — 2023-10-14
- google/timesfm-3.0-pytorch model card (Hugging Face) — 2026-09-25
- ibm-granite/granite-timeseries-ttm-r3 model card (Hugging Face) — 2026-09-25
- ibm-granite/granite-timeseries-patchtst-fm-r2 model card (Hugging Face) — 2026-09-25
- A Time Series is Worth 64 Words: Long-term Forecasting with Transformers — 2022-11-27
- amazon/chronos-2 model card (Hugging Face) — 2026-09-25
- Chronos: Learning the Language of Time Series — 2024-03-12
- autogluon/chronos-2 model card (Hugging Face) — 2026-09-25
- autogluon/chronos-2-small model card (Hugging Face) — 2026-09-25
- Chronos-2: From Univariate to Universal Forecasting — 2025-10-17