1. Language Models (Text → Text)
Candidates grouped by size:
- Small (on-device friendly): Gemma 4 E2B/E4B, Qwen3 0.6B-4B, LFM2.5 1.2B, Nanbeige 4.1 3B
- Mid: Qwen3 8B/14B, Gemma 4 12B, Mistral 7B, Granite 4.0-H
- Large (PCC or Core AI server): Qwen3 MoE, Gemma 4 31B, Qwen3.6 35B, GLM-4.7-Flash
Shared tasks: text classification (RoBERTa), token classification, summarization (T5), Q&A, translation, zero-shot — all viable via FoundationModels @Generable or Core AI LLM.
FoundationModels bridge: Custom LanguageModel provider wraps Core AI LLM → reuse LanguageModelSession API (see running-a-core-ai-model-in-a-foundation-models-session).
2. Audio Pipeline
Microphone → Whisper / Wav2Vec2 (ASR) → text → (Qwen3/Gemma) LLM → Kokoro-82M / VoxCPM-0.5B (TTS) → speaker
↘ Qwen2.5-Omni-3B (audio-text-to-text) for direct audio understanding
↘ CLAP for audio classification
- ASR: Whisper (multilingual robust) vs Wav2Vec2 (streaming) — both in coreai-models.
- TTS: Kokoro 82M (light) vs VoxCPM 0.5B (expressive).
- Audio classification: CLAP; no VAD model in catalog (gap).
3. Vision-Language
Qwen3-VL, MiniCPM-V 4.6, Gemma 4 E2B vision — single model for image-text-to-text, VQA, doc QA, image-to-text. Pair with Unlimited-OCR for document QA; CLIP/EmbeddingGemma/Qwen3-Embedding + Reranker for visual doc retrieval / RAG.
4. Embeddings & Retrieval (RAG)
EmbeddingGemma 300M,Qwen3-Embedding 0.6Bfor sentence similarity / feature extractionQwen3-Reranker 0.6Bfor second-stage ranking- RoBERTa for sentence similarity fallback Use for on-device semantic search before LLM generation.
5. Gaps
Per awesome mapping, no Core AI match for: image-text-to-video, video-text-to-text, any-to-any, unconditional/ video generation, keypoint detection, tabular/time-series. Check Core AI Model Zoo for community fills.