In development

Elix Build AI the Mac-native way

Describe what your app needs to do in Swift. Elix brings composable pipelines, shared resource management, and MLX-based inference for Liquid Foundation Models into one runtime.

  • Composable pipelines
  • Shared service
  • Inference engine
  • On-device

The Elix Stack

Three components, one runtime. Your app describes what it needs to do. Elix handles how it runs.

Your app's request
  1. Composable Pipelines

    1

    Describe the workflow in Swift, using familiar declarative patterns. Elix turns models, tools, adapters, and guardrails into an execution graph and works out what can run in parallel.

  2. Shared Service

    2

    Apps share one AI service instead of loading their own model instances. Elix manages requests and resources across them - and when nobody needs the LLM, it gets out of RAM.

  3. Inference Engine

    3

    Liquid Foundation Models run through an MLX-based engine tuned for Apple silicon, down to custom Metal kernels and per-chip optimizations.

Apple silicon

Composable Pipelines

Define AI pipelines the way you write SwiftUI: declaratively, in native Swift. You describe the intent. Elix handles the orchestration underneath.

A toolchain for AI pipelines

Prompt chains are manageable until they aren't. As they grow, you start wiring dependencies by hand, adapting to model-specific formats, and deciding what can run in parallel. Composable Pipelines turns that wiring into compiler work.

Software
Source code AST Compiler Runtime CPU
Elix
Pipeline DSL Codable AST Compiler Walker Any model

The DSL defines what should happen. Elix handles dependency analysis, parallel scheduling, incremental re-execution, and each model's specific interface underneath.

  • The program is an artifact

    Each pipeline compiles to a Codable AST. It can move between your app and the service, be stored or inspected, and be rendered back as readable pseudo-Swift. The graph is the wire format itself.

  • Schedule by dependency

    The compiler tracks reads and writes, orders dependent steps, and groups independent work into deterministic parallel batches.

  • Models are compile targets

    Chat templates, tool-call formats, and model selection sit below the DSL. Change the model without rewriting the pipeline.

Each step reads the slot the previous one wrote - a straight data-dependency chain.

struct DocumentSummaryPipeline: Pipeline {
    typealias Output = String
    let document: String

    @State var keyPoints = ""
    @State var draft = ""
    @State var summary = ""

    var body: some Pipeline {
        Model<String>("Extract the 5 key points.").message(document).assign(to: $keyPoints)
        Model<String>("Summarize from these points.").input { $keyPoints }.assign(to: $draft)
        Group { Summarize(text: $draft, maxTokens: 512).assign(to: $summary) }
    }
}
Execution-tested examples from the open-source repo - browse the full sources ↗
  • Intent, not infrastructure

    Swift structs with @State and native control flow, composing Model, Guardrail, ForEach, and ClientTask - no YAML, no wiring.

  • Parallel by construction

    The compiler analyses @State reads and writes and batches independent branches automatically - concurrency for free.

  • Compile once, run anywhere

    The Codable graph runs in-process on iOS, across XPC on macOS, or in a test simulator - same pipeline, every target.

Open source - Apache 2.0

The pipeline layer is open source

Composable Pipelines - Elix's authoring, IR, and compiler layer - is available under Apache 2.0. It supports Swift 6.1+ on macOS 14+, iOS 17+, and Linux. An OpenAI-compatible executor is included, so you can run pipelines with Ollama or mlx-lm today.

.package(url: "https://github.com/MacPaw/ComposablePipelines", from: "0.1.0")

Elix as a Shared Service

Instead of every app loading and managing its own LLM, Elix gives them one shared service. Requests, caches, and resources are managed together across clients.

Many clients, one service

App 1App 2EneyApp 4App 5
Shared Service One background process - requests from every client batched, caches shared, and one queue scheduled fairly:
  • Guardrail checks preempt
  • Every app gets its turn
  • Crashed clients reclaimed
One LLM in memory A single resident model serves everyone - loaded on demand, offloaded when clients go inactive.

The memory math

Every app runs its own LLM 15 GB
3 GB3 GB3 GB3 GB3 GB
One shared LLM through Elix 3 GB −80% RAM
3 GB

Five apps share one 3 GB model instead of loading five copies - and when every client goes idle, the service offloads the model entirely: 0 GB.

The privacy boundary

Elix runs the inference. Your app stays in control of files, credentials, APIs, UI, and network access. ClientTask keeps those actions inside the client process, where they belong.

// Pipelines declare intent; ClientTasks own the outside world.
ClientTask(
    input: query,
    action: { encodedQuery in
        // Runs in the client process - full access to
        // keychain, entitlements, UI, file system, network.
        let token  = try Keychain.read("calendar-api-token")
        let result = try await CalendarAPI(token: token).search(encodedQuery)
        return try JSONEncoder().encode(result)
    }
)
  • Everything stays on your Mac

    Prompts, completions, files, secrets, and inference stay local. There's no network path from inference, so Elix works with Wi-Fi off.

  • Telemetry tracks counts, not content

    Only counts and durations leave the Mac over pinned TLS. Your prompts and completions don't.

  • Models are the only download

    Models are downloaded and SHA-256 verified. Once they're on the Mac, Elix is ready to work offline.

  • Only signed clients connect

    Every client is signed and validated before the service accepts its requests.

Purpose-built inference engine

Elix runs Liquid Foundation Models through an MLX-based engine built for Apple silicon. It optimizes each model and tunes it for every chip we support.

How it's built

  • Native from the start

    Built directly on MLX - unified memory, Metal kernels, and pure Swift. No Python runtime overhead between tokens.

  • Focused model support

    Elix supports a focused set of models, so each one gets dedicated optimization - from custom Metal kernels to tuned quantization.

  • Tuned per chip

    Every Apple silicon generation gets its own configuration, selected through measurements on real Macs.

What that gives you

  • Faster first token

    Prefill tuning and model cache reuse reduce the wait before output begins.

  • Faster streaming

    Custom Metal kernels and speculative decoding increase generation throughput.

  • No network round trip

    Inference runs on the Mac, with no server latency, rate limits, or cloud cold starts.

Measured on real hardware

Decode and prefill throughput against the major Apple silicon engines - on real Macs, at prompt lengths from 128 to 32K tokens.

Throughput by prompt length Decode tok/s · higher is better
Average over all prompts tok/s

Decode and prefill measured, median tok/s at each prompt length, ranked by the geometric mean over 128–32K-token contexts. Time to first token is derived as prompt tokens over prefill throughput; end-to-end adds 256 generated tokens at decode throughput. Identical model weights across engines where the format allows: uzu runs its own M-format conversion, llama.cpp runs Q4_0. Each engine was benchmarked through its embedded OpenAI-compatible server, with guidellm as an independent profiling tool. Last updated in September 2026.

Powered by Liquid AI

The model and the runtime work better when they're built with each other in mind. That's the idea behind our work with Liquid AI.

  • Created for Elix

    These aren't off-the-shelf checkpoints. We develop the models around the tasks Mac products need to run.

  • Optimized for the engine

    Liquid Foundation Models are efficient by architecture, then quantized and tuned for Elix's kernels on Apple silicon.

  • Co-developed, shipped together

    Models, inference, and memory are developed as one stack. Eney is the first MacPaw product running on it.

Visit Liquid AI ↗

Beyond text generation

Some jobs just need a fast, focused model. Elix runs compact Liquid AI encoders for tasks like detecting sensitive data, routing prompts, choosing tools, and checking policies. All as steps in the same pipeline.

Multilingual PII detection

Spots 40 kinds of personal data across 16 languages and redacts them in one bidirectional pass - no generation loop, nothing leaves the Mac.

LiquidAI/LFM2.5-Encoder-350M-PII-Detector ↗
Input text

Reschedule my 3pm call with John Smith and text him at +1 234 567 890. Then pay the studio invoice with card 4111 1111 1111 1111 and email the receipt to [email protected] from [email protected].

Detected
  • Identity 1
  • Contact 3
  • Financial 1

What you can build with Elix

Elix brings pipelines, agents, shared inference, and on-device models together for common Mac AI use cases.

Frequently asked questions

The short answers - the rest of the page has the detail.

What is Elix?

Elix is MacPaw's on-device AI runtime for the Mac, currently in development. You describe what your app needs to do in Swift, and Elix handles the rest - composable pipelines, one shared service across apps, and MLX-based inference tuned for Liquid Foundation Models on Apple silicon.

Is Elix open source?

The pipeline layer is. Composable Pipelines - Elix's authoring, IR, and compiler layer - is on GitHub under Apache 2.0, with an OpenAI-compatible executor so you can run pipelines with Ollama or mlx-lm today. The Elix runtime itself is closed source.

Does Elix send my data to the cloud?

Inference runs on the Mac - there is no network path from it, so Elix works with Wi-Fi off. Telemetry carries counts and durations, never prompts or completions. Verified model downloads are the one network hop in the system.

When can I build with Elix?

Elix is in development, and Eney is the first MacPaw product running on the stack. Join the waitlist for early access, or write to [email protected] to talk about a partnership.

Build with Elix

Building something that could use on-device AI? Tell us what you're working on. Join the waitlist for early access to Elix, or get in touch to explore a partnership.

Join the waitlist ↗