Overview
A lightweight, on-device inference engine for FlutterGemma using Dart:ffi, supporting five native platforms and web. It enables efficient local LLM inference via the LiteRT C API embedding backend with opt-in provider integration. Designed for low-latency, privacy-focused generative AI applications directly on devices without relying on cloud services. The package simplifies deployment across Android, iOS, Linux, macOS, Windows, and web environments.
Use cases
- On-device language model inference
- Privacy-preserving chat applications
- Offline AI-powered apps
- Edge computing for LLMs
- Lightweight LLM deployment on mobile and desktop
Key features
- Native platform support (5+ platforms)
- Web compatibility via Dart:ffi
- Opt-in InferenceEngineProvider
- LiteRT C API embedded backend
- Minimal runtime overhead
- Cross-platform consistency
Suitable for
- Developers building offline LLM apps
- Apps requiring data privacy
- Mobile and desktop AI prototypes
- Edge AI projects with limited resources
- Teams using Flutter for generative AI
Considerations
- Requires native code compilation
- Limited to models compatible with LiteRT
- Web support depends on WASM availability
- Performance varies by device hardware
- No built-in model management or caching