Best Forecasting Models for Apple Silicon Macs in 2026: timesfm-3.0-pytorch

Rankings 2026-09-26 Last updated 2026-09-26 12 min read By Q4KM

Quick Answer

Google’s timesfm-3.0-pytorch is the top pick for Apple Silicon Macs as of September 2026 because it is the newest release among the eligible candidates, with no public benchmark yet [1][3][4][6]. In order, the ranking is timesfm-3.0-pytorch (Google), granite-timeseries-patchtst-fm-r2 (IBM Granite), granite-timeseries-ttm-r3 (IBM Granite), and timesfm-2.5-200m-transformers (Google) [1][3].

Key Takeaways

How do these forecasting models compare on parameters, estimated weight memory, context, licenses and benchmarks?

Model Org Params Quant/VRAM Released (date) License Key benchmark (date)
google/timesfm-3.0-pytorch [1] Google 331M [1] 4-bit weights: 165.5 MB estimated (params x 0.5 bytes), from 331M params [1]; runtime VRAM not published 2026-08-24 [1] timesfm-non-commercial-license-v1.0 [1] no public benchmark yet
ibm-granite/granite-timeseries-patchtst-fm-r2 [4] IBM Granite 385M [4] 4-bit weights: 192.5 MB estimated (params x 0.5 bytes), from 385M params [4]; runtime VRAM not published 2026-08-07 [4] openmdw-1.0 [4] no public benchmark yet
ibm-granite/granite-timeseries-ttm-r3 [6] IBM Granite 1M [6] 4-bit weights: 0.5 MB estimated (params x 0.5 bytes), from 1M params [6]; runtime VRAM not published 2026-05-21 [6] Apache-2.0 [6] no public benchmark yet
google/timesfm-2.5-200m-transformers [3] Google 231M [3] 4-bit weights: 115.5 MB estimated (params x 0.5 bytes), from 231M params [3]; runtime VRAM not published 2026-02-19 [3] Apache-2.0 [3] no public benchmark yet

Which forecasting models should you consider for an Apple Silicon Mac?

1. timesfm-3.0-pytorch

timesfm-3.0-pytorch by Google ranks first because its Hugging Face publication date of August 24, 2026, is the newest among the eligible releases.[1][3][4][6] The model has 331M parameters and a published F32 weight footprint of 1.3 GB.[1] Its forecasting family is associated with Google’s paper “A decoder-only foundation model for time-series forecasting.”[2] Benchmark status: no public benchmark yet; the placement does not establish a forecasting accuracy advantage.

For local planning, its 331M parameters imply a hypothetical 4-bit weight footprint of 165.5 MB, estimated (params x 0.5 bytes).[1] Treat that as a weight-storage calculation, not a total memory requirement or confirmation of quantized runtime support. A specific Apple Silicon Mac configuration cannot be recommended from those figures alone. Before choosing hardware, verify runtime compatibility, supported context length and total memory use. Consider the model for non-commercial forecasting evaluation; its timesfm-non-commercial-license-v1.0 license is the key caveat for commercial deployment.[1]

Ranking note: this ranking admits only time series forecasting models from labs with a published paper or leaderboard record (Hugging Face pipeline tags: time-series-forecasting). With no public benchmark scoring any candidate, the ordering is newest release first, then downloads.[1][3][4][6] Popularity serves only as a tiebreak; parameter count and estimated memory do not determine placement.

2. granite-timeseries-patchtst-fm-r2

granite-timeseries-patchtst-fm-r2 by IBM Granite ranks second because its publication date falls behind Google’s TimesFM release and ahead of the remaining candidates.[1][3][4][6] The ordering is newest release first, then downloads; no public benchmark scores any candidate.[1][3][4][6] The ranking admits only time series forecasting models from labs with a published paper or leaderboard record, using the Hugging Face pipeline tag time-series-forecasting.

Published on August 7, 2026, the model has 385M parameters, F32 weights occupying 1.5 GB, and the openmdw-1.0 license.[4] Weight storage at 4-bit is 192.5 MB, estimated (params x 0.5 bytes) from the cited 385M parameters.[4] For an Apple Silicon Mac, use that estimate as a starting point for hardware planning. A total RAM requirement, context limit, and verified Apple GPU execution path remain unspecified.

Long-term forecasting is a practical evaluation target, consistent with the associated paper, “A Time Series is Worth 64 Words: Long-term Forecasting with Transformers.”[5] Consider the model for local experiments when its license suits your project.[4] The caveat is straightforward: no public benchmark yet establishes its comparative forecasting accuracy. Its position reflects release order, so evaluate forecast quality and total runtime memory on your intended workload before committing to a Mac configuration.

3. granite-timeseries-ttm-r3

granite-timeseries-ttm-r3 from IBM Granite ranks third because its first Hugging Face publication falls after the Google TimesFM and IBM Granite PatchTST entries ahead of it and before the remaining Google entry.[1][3][4][6] The ordering basis is newest release first, then downloads; no public benchmark scores any candidate. The ranking admits only time-series forecasting models from labs with a published paper or leaderboard record, using the Hugging Face pipeline tag time-series-forecasting.

IBM Granite lists 1M parameters, an Apache-2.0 license and a first publication date of May 21, 2026.[6] Weight storage at 4-bit is approximately 0.5 MB, estimated (params x 0.5 bytes) from the cited 1M parameters.[6] Treat that calculation as a weight-storage budget when planning a local installation. A specific Apple Silicon processor or unified-memory requirement cannot be established from the published specifications, and the weight estimate does not establish total runtime memory or Mac compatibility.[6]

Consider granite-timeseries-ttm-r3 for a local forecasting evaluation when a compact parameter count and Apache-2.0 licensing match your project requirements.[6] The benchmark status is no public benchmark yet: its position does not establish forecast accuracy or execution speed relative to the other candidates. Before committing to a Mac deployment, check the supported execution path and evaluate forecasts on your own series. Context length and measured Apple Silicon performance remain unspecified.[6]

4. timesfm-2.5-200m-transformers

timesfm-2.5-200m-transformers from Google ranks fourth because its February 19, 2026 publication precedes the other candidates’ releases.[3][1][4][6] The ordering basis is “newest release first, then downloads”; no public benchmark yet establishes a performance ranking among the candidates. The ranking admits only time series forecasting models from labs with a published paper or leaderboard record, using Hugging Face’s time-series-forecasting pipeline tag. Its position therefore reflects release timing, not demonstrated forecasting accuracy.

The model has 231M parameters, a context length of 16K tokens, and an F32 weight download of 0.9 GB.[3] The license is Apache-2.0.[3] Consider it for local forecasting experiments where the documented context window and Apache licensing match your requirements.[3] Google’s model family is associated with the paper “A decoder-only foundation model for time-series forecasting.”[2] No published comparative score here establishes an accuracy advantage for this checkpoint.

For an Apple Silicon Mac, quantized weight storage would be 115.5 MB—estimated (params x 0.5 bytes), using the cited 231M parameters.[3] Treat that calculation as a planning figure for weights alone. The caveat is that weight size does not establish a working quantized runtime, Apple GPU support, or total unified-memory requirements. A specific Mac configuration cannot be recommended from the documented specifications alone; validate runtime compatibility and actual memory use before committing hardware.

How do you estimate quantized weight memory for local forecasting?

Estimate local forecasting weight memory by multiplying the published parameter count by 0.5 bytes per parameter for 4-bit storage, and label the result “estimated (params x 0.5 bytes).”[1][3][4][6] Use the parameter count rather than inferring it from a model’s name, and keep weight storage separate from the memory needed to run a forecast.

Google’s timesfm-3.0-pytorch (Google) has 331M parameters: approximately 165.5 MB of weight storage, estimated (params x 0.5 bytes).[1] Google’s timesfm-2.5-200m-transformers (Google) has 231M parameters: approximately 115.5 MB, estimated (params x 0.5 bytes).[3]

IBM Granite’s ibm-granite/granite-timeseries-patchtst-fm-r2 has 385M parameters: approximately 192.5 MB, estimated (params x 0.5 bytes).[4] IBM Granite’s ibm-granite/granite-timeseries-ttm-r3 has 1M parameters: approximately 0.5 MB, estimated (params x 0.5 bytes).[6] Treat each calculation as a planning estimate, rather than the size of an available quantized download.

Check the published weight format before comparing those estimates with repository sizes. Google lists F32 weights of 1.3 GB for timesfm-3.0-pytorch (Google), rather than a quantized package matching the estimate above.[1] A weight calculation alone does not establish Apple Silicon compatibility, total runtime memory, or whether a particular Mac can run the model. Report those separately when measured; do not turn an estimated weight size into a RAM recommendation.

Which licenses apply to these forecasting models?

The applicable licenses are timesfm-non-commercial-license-v1.0 for Google’s timesfm-3.0-pytorch (Google) [1], openmdw-1.0 for IBM Granite’s ibm-granite/granite-timeseries-patchtst-fm-r2 [4], and Apache-2.0 for IBM Granite’s ibm-granite/granite-timeseries-ttm-r3 [6] and Google’s timesfm-2.5-200m-transformers (Google) [3].

Google’s TimesFM releases therefore require separate license checks: the listed PyTorch release carries the TimesFM non-commercial license [1], while the listed Transformers release carries Apache-2.0 [3]. For an application already using TimesFM, treat a change of model repository as a reason to review licensing again. Do not assume that a shared family name means shared terms.

IBM Granite also uses different licenses across the listed forecasting models. The PatchTST release carries openmdw-1.0 [4]; the Tiny Time Mixer release carries Apache-2.0 [6]. Choose the exact model before completing a license review, and record its repository identifier alongside the applicable license.

For a local Mac deployment, review the selected license against your intended use, including any commercial application, modification or redistribution. Keep that review separate from memory and runtime evaluation. An available weight download should prompt a license check before adoption; it should not substitute for one.

What remains unverified about forecasting performance on Apple Silicon?

Forecasting accuracy, inference speed, runtime memory use and execution support on Apple Silicon remain unverified for the candidates, so their release order does not establish a Mac performance winner.

Google’s timesfm-3.0-pytorch (Google) has no public benchmark yet [1]. IBM Granite’s ibm-granite/granite-timeseries-patchtst-fm-r2 has no public benchmark yet [4], and its ibm-granite/granite-timeseries-ttm-r3 has no public benchmark yet [6]. Google’s timesfm-2.5-200m-transformers (Google) also has no public benchmark yet [3]. Forecast quality therefore remains an open comparison; publication dates and download counts cannot resolve it.

Published parameter counts and stored weight sizes do not establish runtime memory requirements [1][3][4][6]. A practical Mac evaluation still needs to measure memory during loading and forecasting, confirm the execution backend, and test whether quantization preserves useful predictions. A calculated weight footprint alone would not demonstrate that a model fits or runs well on a particular Mac.

Google’s timesfm-2.5-200m-transformers (Google) lists a context length of 16K tokens [3]. Usable context under Mac memory constraints remains unverified, as does the effect of longer inputs on latency and forecast quality. Comparable context limits are not established here for the other candidates.

A useful local comparison would hold datasets, forecast horizons, input lengths and evaluation metrics constant, while recording hardware, software, precision, latency and peak memory. Until those measurements exist, treat the ranking as a release-based shortlist for evaluation.

Frequently Asked Questions

Which forecasting model should I evaluate first on an Apple Silicon Mac?

Google’s timesfm-3.0-pytorch (Google) ranks first because its 2026-08-24 release is the newest eligible release.[1][3][4][6] The model has 331M parameters and uses timesfm-non-commercial-license-v1.0.[1] Its benchmark status is no public benchmark yet.[1] Treat this position as an evaluation order, rather than evidence of forecasting accuracy or measured Apple Silicon performance.

How are the forecasting models ranked?

Scope: this ranking admits only time series forecasting models from labs with a published paper or leaderboard record (Hugging Face pipeline tags: time-series-forecasting). Ranking note: no public benchmark scores any candidate, so the rule gives: newest release first, then downloads.[1][3][4][6] The order is Google’s timesfm-3.0-pytorch (Google),[1] IBM’s (ibm-granite) ibm-granite/granite-timeseries-patchtst-fm-r2,[4] IBM’s ibm-granite/granite-timeseries-ttm-r3,[6] then Google’s timesfm-2.5-200m-transformers (Google).[3]

How much memory would quantized weights need?

For 4-bit weights, timesfm-3.0-pytorch (Google) has 331M parameters: 165.5 MB estimated (params x 0.5 bytes).[1] ibm-granite/granite-timeseries-patchtst-fm-r2 has 385M parameters: 192.5 MB estimated (params x 0.5 bytes).[4] ibm-granite/granite-timeseries-ttm-r3 has 1M parameters: 0.5 MB estimated (params x 0.5 bytes).[6] timesfm-2.5-200m-transformers (Google) has 231M parameters: 115.5 MB estimated (params x 0.5 bytes).[3] Those calculations cover weights alone; they do not establish total runtime memory or availability of a working quantized implementation.

Which model specifies a context length for planning input history?

Google’s timesfm-2.5-200m-transformers (Google) specifies a context length of 16K tokens.[3] Use that limit when evaluating whether the model suits your forecasting workflow. Check how the implementation represents observations before translating tokens into a history window. A context specification alone does not establish forecasting accuracy, inference speed or practical memory use on an Apple Silicon Mac.

Which licenses do these forecasting models use?

Google’s timesfm-3.0-pytorch (Google) uses timesfm-non-commercial-license-v1.0.[1] IBM’s ibm-granite/granite-timeseries-patchtst-fm-r2 uses openmdw-1.0.[4] IBM’s ibm-granite/granite-timeseries-ttm-r3 and Google’s timesfm-2.5-200m-transformers (Google) use Apache-2.0.[6][3] Check the applicable license before choosing a model for a product or internal deployment. In particular, do not assume that different releases within the Google TimesFM family share the same license.

Can I compare forecasting accuracy from published benchmark scores?

The benchmark status for timesfm-3.0-pytorch (Google) is no public benchmark yet;[1] for ibm-granite/granite-timeseries-patchtst-fm-r2, no public benchmark yet;[4] for ibm-granite/granite-timeseries-ttm-r3, no public benchmark yet;[6] and for timesfm-2.5-200m-transformers (Google), no public benchmark yet.[3] Release dates and download counts therefore should not be read as an accuracy comparison. Evaluate candidates against your own forecasting task before selecting one.

Sources

  1. google/timesfm-3.0-pytorch model card (Hugging Face) — 2026-09-25
  2. A decoder-only foundation model for time-series forecasting — 2023-10-14
  3. google/timesfm-2.5-200m-transformers model card (Hugging Face) — 2026-09-25
  4. ibm-granite/granite-timeseries-patchtst-fm-r2 model card (Hugging Face) — 2026-09-25
  5. A Time Series is Worth 64 Words: Long-term Forecasting with Transformers — 2022-11-27
  6. ibm-granite/granite-timeseries-ttm-r3 model card (Hugging Face) — 2026-09-25

Get these models on a hard drive

Skip the downloads. Browse our catalog of 985+ commercially-licensed AI models, available pre-loaded on high-speed drives.

Browse Model Catalog