Run local LLMs like Gemma, Qwen, and LLaMA on Android for offline, private, real-time chat and question answering with LiteRT and ONNX Runtime.
-
Updated
Sep 26, 2026 - Kotlin
Run local LLMs like Gemma, Qwen, and LLaMA on Android for offline, private, real-time chat and question answering with LiteRT and ONNX Runtime.
Training pipeline for an osu!mania 7k next-event model, designed as the upstream predictor for audio-driven 7k map generation systems.
离线可用的本地类型化决策:4 核 CPU 单题 15.6ms。Local & offline Jev / System One inference on CPU — ONNX + INT8, no torch at runtime. 支持 laya / kev / PlayJev
把本地 System-1 决策模型(Laya)包成 HTTP 服务:意图分类 / 工单分派 / LLM 路由 / 内容审核,离线零 token、不生成文字;Apple Silicon(MLX) 与 Linux(torch) 双后端;可导出 OpenAPI 3.0/3.1;Java/TS/Python 零依赖 SDK。
A Cloud-to-Edge MLOps pipeline for offline industrial diagnostics. Fine-tunes Phi-3-mini (3.8B) on Cloud GPUs via QLoRA, quantizes to INT4, and deploys as a CPU-optimized ONNX microservice for industrial standard sensor logs.
이미지·상황·질문에서 정답을 고르되 근거가 부족하면 '모름'을 택하도록 학습시킨 멀티모달 편향 QA — 성균관대 챌린지
A comprehensive toolkit for streamlining and simplifying the offline inference process for LLMs across various models and libraries.
Download Hugging Face model snapshots locally for offline inference, deployment, and Docker workflows.
Offline Android flower identification using Kotlin, Jetpack Compose, and a 102-class TensorFlow Lite MobileNetV2 model.
Event-driven image processing pipeline with automatic format normalization, thumbnail generation, AI-powered captioning (BLIP), and full lifecycle management. Built with .NET 10, Python, Docker, and Terraform.
Offline Krita plugin for e621 image tag generation using JTP-3 ONNX Runtime CPU inference
Offline satellite image-tile land-use classifier — a FastAPI service running a local ONNX model on CPU, storing predictions in SQLite for analysts to query. Built to run air-gapped.
Real-time semantic audio codec achieving 300bps bandwidth via gen AI reconstruction.
Aptus-R: An offline candidate ranking system using dual retrieval (FAISS + BM25), a 5-signal composite, and local Phi-3-mini reranking.
Inclusive hand gesture recognition system for assistive human–computer interaction, based on classical machine learning and MediaPipe Hands.
GPT-OSS B20 Local Execution. Lightweight local environment for running it with Python 3.12 and CUDA acceleration. - Run GPT-OSS B20 entirely offline - Optimize text generation with GPU - Enable fast, secure inference on consumer hardware.
Мультимодальная офлайновая система детекции контрафакта (текст+изображение+таблица)
Offline CrowdAware system for Raspberry Pi 4B and Heltec LoRa V3 using Raspberry Pi Camera Module 3 and MLX90640 Thermal Camera.
To associate your repository with the offline-inference topic, visit your repo's landing page and select "manage topics."