---
title: "Untitled"
url: https://memory.wiki/xYoo54CC
updated: 2026-10-09T20:44:23.290Z
hub: https://memory.wiki/hub/pratofeito
concept_count: 11
source: "desktop"
---
---

title: "Language, Audio & Multimodal Pipelines — Extensive" description: "LLMs (Qwen/Gemma/Mistral), embeddings, ASR/TTS and any-to-any gaps." type: article status: stable tags: \[core-ai, llm, audio, extensive\] sources:

- resource: ../external-sources/awesome-core-ai-models.md
- resource: ../external-sources/apple-coreai-models.md

---

## 1. Language Models (Text → Text)

Candidates grouped by size:

- **Small (on-device friendly):** Gemma 4 E2B/E4B, Qwen3 0.6B-4B, LFM2.5 1.2B, Nanbeige 4.1 3B
- **Mid:** Qwen3 8B/14B, Gemma 4 12B, Mistral 7B, Granite 4.0-H
- **Large (PCC or Core AI server):** Qwen3 MoE, Gemma 4 31B, Qwen3.6 35B, GLM-4.7-Flash

Shared tasks: text classification (RoBERTa), token classification, summarization (T5), Q&A, translation, zero-shot — all viable via FoundationModels `@Generable` or Core AI LLM.

FoundationModels bridge: `Custom LanguageModel` provider wraps Core AI LLM → reuse `LanguageModelSession` API (see `running-a-core-ai-model-in-a-foundation-models-session`).

## 2. Audio Pipeline

```text
Microphone → Whisper / Wav2Vec2 (ASR) → text → (Qwen3/Gemma) LLM → Kokoro-82M / VoxCPM-0.5B (TTS) → speaker
        ↘ Qwen2.5-Omni-3B (audio-text-to-text) for direct audio understanding
        ↘ CLAP for audio classification
```

- **ASR:** Whisper (multilingual robust) vs Wav2Vec2 (streaming) — both in coreai-models.
- **TTS:** Kokoro 82M (light) vs VoxCPM 0.5B (expressive).
- **Audio classification:** CLAP; no VAD model in catalog (gap).

## 3. Vision-Language

Qwen3-VL, MiniCPM-V 4.6, Gemma 4 E2B vision — single model for image-text-to-text, VQA, doc QA, image-to-text. Pair with Unlimited-OCR for document QA; CLIP/EmbeddingGemma/Qwen3-Embedding + Reranker for visual doc retrieval / RAG.

## 4. Embeddings & Retrieval (RAG)

- `EmbeddingGemma 300M`, `Qwen3-Embedding 0.6B` for sentence similarity / feature extraction
- `Qwen3-Reranker 0.6B` for second-stage ranking
- RoBERTa for sentence similarity fallback Use for on-device semantic search before LLM generation.

## 5. Gaps

Per awesome mapping, no Core AI match for: image-text-to-video, video-text-to-text, any-to-any, unconditional/ video generation, keypoint detection, tabular/time-series. Check Core AI Model Zoo for community fills.

See [Catalog](./core-ai-model-catalog.md) · [Bindings](./apple-on-device-ai-bindings.md) · [Vision](./core-ai-vision-segmentation.md)

---

## Summary
The document outlines a framework for integrating language, audio, and vision models into pipelines using specific on-device and server-side AI tools. While these models support tasks like text generation, speech processing, and retrieval, the current catalog lacks capabilities for video generation, time-series analysis, and any-to-any model conversion.

## Themes
- AI model pipelines
- On-device model architecture
- Multimodal integration strategies

## Key takeaways
- Language models are categorized into small, mid, and large tiers to support different deployment environments.
- FoundationModels provides a bridge to reuse the LanguageModelSession API for custom model providers.
- The audio pipeline utilizes Whisper or Wav2Vec2 for ASR and Kokoro or VoxCPM for TTS.
- Qwen3-VL and MiniCPM-V are recommended for tasks including VQA, document QA, and image to text.
- EmbeddingGemma and Qwen3-Embedding are the primary models used for sentence similarity and feature extraction.
- The document explicitly identifies a lack of Core AI matches for video generation, keypoint detection, and tabular or time-series analysis.

## Insights
- The architecture prioritizes a tiered approach to language models based on hardware constraints, ranging from on-device small models to server-side large models.
- The audio pipeline demonstrates a modular design where specialized models for ASR and TTS are bridged by LLMs to create conversational interfaces.
- Visual document retrieval relies on a combination of vision-language models and dedicated reranking models to improve RAG performance.

## Open questions / gaps
- What specific community models are available in the Core AI Model Zoo to address the identified gaps in video and tabular data processing?

## Concepts in this document
- **RAG** _(concept)_
  Retrieval-augmented generation using embeddings and rerankers.
- **Language Models** _(concept)_
  Categorization of LLMs by size for text-based tasks.
- **Core AI** _(tag)_
  The primary infrastructure and model ecosystem described.
- **Audio Pipeline** _(concept)_
  Workflow integrating ASR, LLMs, and TTS for audio processing.
- **Qwen3** _(entity)_
  A versatile model family used across text, vision, and embedding tasks.
- **Gaps** _(concept)_
  Identified missing capabilities in the current Core AI model catalog.
- **Gemma 4** _(entity)_
  A model family utilized for text and vision-language capabilities.
- **Vision-Language** _(concept)_
  Models for image-text interaction and document QA.
- **Kokoro-82M** _(entity)_
  Lightweight TTS model for audio output.
- **CLAP** _(entity)_
  Model used for audio classification tasks.
- **Whisper** _(entity)_
  ASR model used for multilingual speech recognition.

## Concept relations (within this doc's concepts)
- **Language Models** includes model family **Gemma 4**
- **Audio Pipeline** uses for tts **Kokoro-82M**
- **RAG** uses for embedding **Qwen3**
- **Core AI** has missing features **Gaps**
- **Language Models** includes model family **Qwen3**
- **Audio Pipeline** uses for asr **Whisper**
- **Audio Pipeline** uses for classification **CLAP**
- **Vision-Language** uses for vision **Gemma 4**

_Hub canonical:_ https://memory.wiki/hub/pratofeito
_Concept digest:_ https://memory.wiki/raw/hub/pratofeito?digest=1&compact=1
