1. Language Models (Text → Text)

Candidates grouped by size:

  • Small (on-device friendly): Gemma 4 E2B/E4B, Qwen3 0.6B-4B, LFM2.5 1.2B, Nanbeige 4.1 3B
  • Mid: Qwen3 8B/14B, Gemma 4 12B, Mistral 7B, Granite 4.0-H
  • Large (PCC or Core AI server): Qwen3 MoE, Gemma 4 31B, Qwen3.6 35B, GLM-4.7-Flash

Shared tasks: text classification (RoBERTa), token classification, summarization (T5), Q&A, translation, zero-shot — all viable via FoundationModels @Generable or Core AI LLM.

FoundationModels bridge: Custom LanguageModel provider wraps Core AI LLM → reuse LanguageModelSession API (see running-a-core-ai-model-in-a-foundation-models-session).

2. Audio Pipeline

Microphone → Whisper / Wav2Vec2 (ASR) → text → (Qwen3/Gemma) LLM → Kokoro-82M / VoxCPM-0.5B (TTS) → speaker ↘ Qwen2.5-Omni-3B (audio-text-to-text) for direct audio understanding ↘ CLAP for audio classification
  • ASR: Whisper (multilingual robust) vs Wav2Vec2 (streaming) — both in coreai-models.
  • TTS: Kokoro 82M (light) vs VoxCPM 0.5B (expressive).
  • Audio classification: CLAP; no VAD model in catalog (gap).

3. Vision-Language

Qwen3-VL, MiniCPM-V 4.6, Gemma 4 E2B vision — single model for image-text-to-text, VQA, doc QA, image-to-text. Pair with Unlimited-OCR for document QA; CLIP/EmbeddingGemma/Qwen3-Embedding + Reranker for visual doc retrieval / RAG.

4. Embeddings & Retrieval (RAG)

  • EmbeddingGemma 300M, Qwen3-Embedding 0.6B for sentence similarity / feature extraction
  • Qwen3-Reranker 0.6B for second-stage ranking
  • RoBERTa for sentence similarity fallback Use for on-device semantic search before LLM generation.

5. Gaps

Per awesome mapping, no Core AI match for: image-text-to-video, video-text-to-text, any-to-any, unconditional/ video generation, keypoint detection, tabular/time-series. Check Core AI Model Zoo for community fills.

See Catalog · Bindings · Vision