Back to news
Large Language Models
2d ago

Cactus releases Needle 2, a compact LLM for various devices

Aug 10, 2026
AI Summary

Cactus has launched Needle 2, a 14MB agentic language model designed for phones, wearables, and smart home devices. The model operates efficiently on low-resource hardware, achieving high decoding speeds while consuming significantly less power compared to traditional models.

  • Needle 2 is a 14MB language model designed for use in mobile devices, wearables, smart homes, and small robots.
  • It operates with 45 million parameters at 2-bit compression and requires only 28MB of RAM for a full session.
  • The model achieves decoding speeds of 500 tokens per second on Raspberry Pi 5, and between 300-1,500 tokens per second on various VR devices and budget smartphones.
  • Needle 2 is designed to perform structured extraction and can be fine-tuned for specific tasks using a Python package.
  • The model is significantly smaller and more efficient than comparable models, consuming 7x to 85x fewer resources per token.
  • It incorporates a confidence scoring system to determine when to act or escalate tasks to more powerful models.
  • Users are encouraged to test Needle 2 and provide feedback to improve its functionality.
llmagenticsmart homewearablesraspberry pi