Files
timesfm/timesfm-forecasting/references/system_requirements.md
T
borealBytes 5aad77bd61 docs: address PR #369 review comments and add dataset preflight
- Add context limit rationale to system_requirements.md with memory formula
- Update SKILL.md to include XReg/covariates in description and usage sections
- Add dataset-aware memory estimation to check_system.py with new CLI args
- Document memory estimation in api_reference.md with Mermaid diagram
- Add dataset preflight section to SKILL.md with examples

Resolves review comments about:
- How context limits (512/1024) were determined
- Including XReg mode description in skill documentation

Bonus enhancement: Dataset preflight checking prevents OOM before loading data.
2026-02-25 21:27:31 -05:00

6.9 KiB
Raw Blame History

System Requirements for TimesFM

Hardware Tiers

TimesFM can run on a variety of hardware configurations. This guide helps you choose the right setup and tune performance for your machine.

How Context Limits Are Determined

The max_context values in each tier are conservative recommendations based on memory-performance tradeoffs, not hard limits. TimesFM 2.5 supports up to 16,384 context points, but smaller values are recommended for most use cases.

Why 512 and 1024?

Factor 512 Context 1024 Context
Memory per 1000 series ~100 MB ~200 MB
Typical Use Case Daily data, ~1-2 years Daily data, ~2-3 years
Inference Speed Faster Moderate
Hardware 4-8 GB RAM 16 GB RAM or GPU

Memory Formula: RAM ≈ model_weights + 0.5 GB + (0.2 MB × num_series × context_length / 1000)

Where:

  • model_weights = ~800 MB (TimesFM 2.5)
  • context_length = your max_context value
  • num_series = number of time series in your batch

You can use larger contexts if your hardware supports it:

  • Up to 2048: Requires ~16 GB RAM for moderate batch sizes
  • Up to 4096: Requires GPU or 32+ GB RAM
  • Up to 16384: Maximum supported, requires significant memory

See Data Preparation Guide for context length recommendations by data frequency.

Tier 1: Minimal (CPU-Only, 48 GB RAM)

  • Use case: Light exploration, single-series forecasting, prototyping
  • Model: TimesFM 2.5 (200M) only
  • Batch size: per_core_batch_size=4
  • Context: Limit max_context=512
  • Expected speed: ~25 seconds per 100-point series
model.compile(timesfm.ForecastConfig(
    max_context=512,
    max_horizon=128,
    per_core_batch_size=4,
    normalize_inputs=True,
    use_continuous_quantile_head=True,
    fix_quantile_crossing=True,
))

Tier 2: Standard (CPU 16 GB or GPU 48 GB VRAM)

  • Use case: Batch forecasting (dozens of series), evaluation, production prototypes
  • Model: TimesFM 2.5 (200M)
  • Batch size: per_core_batch_size=32 (CPU) or 64 (GPU)
  • Context: max_context=1024
  • Expected speed: ~0.51 second per 100-point series (GPU)
model.compile(timesfm.ForecastConfig(
    max_context=1024,
    max_horizon=256,
    per_core_batch_size=64,
    normalize_inputs=True,
    use_continuous_quantile_head=True,
    fix_quantile_crossing=True,
))

Tier 3: Production (GPU 16+ GB VRAM or Apple Silicon 32+ GB)

  • Use case: Large-scale batch forecasting (thousands of series), long context
  • Model: TimesFM 2.5 (200M)
  • Batch size: per_core_batch_size=128256
  • Context: max_context=4096 or higher
  • Expected speed: ~0.10.3 seconds per 100-point series
model.compile(timesfm.ForecastConfig(
    max_context=4096,
    max_horizon=256,
    per_core_batch_size=128,
    normalize_inputs=True,
    use_continuous_quantile_head=True,
    fix_quantile_crossing=True,
))

Tier 4: Legacy Models (v1.0/v2.0 — 500M parameters)

  • ⚠️ WARNING: TimesFM v2.0 (500M) requires ≥ 16 GB RAM (CPU) or ≥ 8 GB VRAM (GPU)
  • ⚠️ WARNING: TimesFM v1.0 legacy JAX version may require ≥ 32 GB RAM
  • Recommendation: Unless you specifically need a legacy checkpoint, use TimesFM 2.5

Memory Estimation

CPU Memory (RAM)

Approximate RAM usage during inference:

Component TimesFM 2.5 (200M) TimesFM 2.0 (500M)
Model weights ~800 MB ~2 GB
Runtime overhead ~500 MB ~1 GB
Input/output buffers ~200 MB per 1000 series ~500 MB per 1000 series
Total (small batch) ~1.5 GB ~3.5 GB
Total (large batch) ~3 GB ~6 GB

Formula: RAM ≈ model_weights + 0.5 GB + (0.2 MB × num_series × context_length / 1000)

GPU Memory (VRAM)

Component TimesFM 2.5 (200M)
Model weights ~800 MB
KV cache + activations ~200500 MB (scales with context)
Batch buffers ~100 MB per 100 series at context=1024
Total (batch=32) ~1.2 GB
Total (batch=128) ~1.8 GB
Total (batch=256) ~2.5 GB

Disk Space

Item Size
TimesFM 2.5 safetensors ~800 MB
Hugging Face cache overhead ~200 MB
Total download ~1 GB

Model weights are downloaded once from Hugging Face Hub and cached in ~/.cache/huggingface/ (or $HF_HOME).

GPU Selection Guide

NVIDIA GPUs (CUDA)

GPU VRAM Recommended batch Notes
RTX 3060 12 GB 64 Good entry-level
RTX 3090 / 4090 24 GB 256 Excellent for production
A100 (40 GB) 40 GB 512 Cloud/HPC
A100 (80 GB) 80 GB 1024 Cloud/HPC
T4 16 GB 128 Cloud (Colab, AWS)
V100 1632 GB 128256 Cloud

Apple Silicon (MPS)

Chip Unified Memory Recommended batch Notes
M1 816 GB 1632 Works, slower than CUDA
M1 Pro/Max 1664 GB 32128 Good performance
M2/M3/M4 Pro/Max 18128 GB 64256 Excellent

CPU Only

Works on any CPU with sufficient RAM. Expect 520× slower than GPU.

Python and Package Requirements

Requirement Minimum Recommended
Python 3.10 3.12+
numpy 1.26.4 latest
torch 2.0.0 latest
huggingface_hub 0.23.0 latest
safetensors 0.5.3 latest

Optional Dependencies

Package Purpose Install
jax Flax backend pip install jax[cuda]
flax Flax backend pip install flax
scikit-learn XReg covariates pip install scikit-learn

Operating System Compatibility

OS Status Notes
Linux (Ubuntu 20.04+) Fully supported Best performance with CUDA
macOS 13+ (Ventura) Fully supported MPS acceleration on Apple Silicon
Windows 11 + WSL2 Supported Use WSL2 for best experience
Windows (native) ⚠️ Partial PyTorch works, some edge cases

Troubleshooting

Out of Memory (OOM)

# Reduce batch size
model.compile(timesfm.ForecastConfig(
    per_core_batch_size=4,  # Start very small
    max_context=512,        # Reduce context
    ...
))

# Process in chunks
for i in range(0, len(inputs), 50):
    chunk = inputs[i:i+50]
    p, q = model.forecast(horizon=H, inputs=chunk)

Slow Inference on CPU

# Ensure matmul precision is set
import torch
torch.set_float32_matmul_precision("high")

# Use smaller context
model.compile(timesfm.ForecastConfig(
    max_context=256,  # Shorter context = faster
    ...
))

Model Download Fails

# Set a different cache directory
export HF_HOME=/path/with/more/space

# Or download manually
huggingface-cli download google/timesfm-2.5-200m-pytorch