Describe what your app needs to do in Swift. Elix brings composable pipelines, shared resource management, and MLX-based inference for Liquid Foundation Models into one runtime.
Three components, one runtime. Your app describes what it needs to do. Elix handles how it runs.
Describe the workflow in Swift, using familiar declarative patterns. Elix turns models, tools, adapters, and guardrails into an execution graph and works out what can run in parallel.
Apps share one AI service instead of loading their own model instances. Elix manages requests and resources across them - and when nobody needs the LLM, it gets out of RAM.
Liquid Foundation Models run through an MLX-based engine tuned for Apple silicon, down to custom Metal kernels and per-chip optimizations.
Define AI pipelines the way you write SwiftUI: declaratively, in native Swift. You describe the intent. Elix handles the orchestration underneath.
Prompt chains are manageable until they aren't. As they grow, you start wiring dependencies by hand, adapting to model-specific formats, and deciding what can run in parallel. Composable Pipelines turns that wiring into compiler work.
The DSL defines what should happen. Elix handles dependency analysis, parallel scheduling, incremental re-execution, and each model's specific interface underneath.
Each pipeline compiles to a Codable AST. It can move between your app and the service, be stored or inspected, and be rendered back as readable pseudo-Swift. The graph is the wire format itself.
The compiler tracks reads and writes, orders dependent steps, and groups independent work into deterministic parallel batches.
Chat templates, tool-call formats, and model selection sit below the DSL. Change the model without rewriting the pipeline.
Each step reads the slot the previous one wrote - a straight data-dependency chain.
struct DocumentSummaryPipeline: Pipeline {
typealias Output = String
let document: String
@State var keyPoints = ""
@State var draft = ""
@State var summary = ""
var body: some Pipeline {
Model<String>("Extract the 5 key points.").message(document).assign(to: $keyPoints)
Model<String>("Summarize from these points.").input { $keyPoints }.assign(to: $draft)
Group { Summarize(text: $draft, maxTokens: 512).assign(to: $summary) }
}
} Execution-tested examples from the open-source repo - browse the full sources ↗Swift structs with @State and native control flow, composing Model, Guardrail, ForEach, and ClientTask - no YAML, no wiring.
The compiler analyses @State reads and writes and batches independent branches automatically - concurrency for free.
The Codable graph runs in-process on iOS, across XPC on macOS, or in a test simulator - same pipeline, every target.
Open source - Apache 2.0
Composable Pipelines - Elix's authoring, IR, and compiler layer - is available under Apache 2.0. It supports Swift 6.1+ on macOS 14+, iOS 17+, and Linux. An OpenAI-compatible executor is included, so you can run pipelines with Ollama or mlx-lm today.
.package(url: "https://github.com/MacPaw/ComposablePipelines", from: "0.1.0") Instead of every app loading and managing its own LLM, Elix gives them one shared service. Requests, caches, and resources are managed together across clients.
Many clients, one service
The memory math
Five apps share one 3 GB model instead of loading five copies - and when every client goes idle, the service offloads the model entirely: 0 GB.
Elix runs the inference. Your app stays in control of files, credentials, APIs, UI, and network access. ClientTask keeps those actions inside the client process, where they belong.
// Pipelines declare intent; ClientTasks own the outside world.
ClientTask(
input: query,
action: { encodedQuery in
// Runs in the client process - full access to
// keychain, entitlements, UI, file system, network.
let token = try Keychain.read("calendar-api-token")
let result = try await CalendarAPI(token: token).search(encodedQuery)
return try JSONEncoder().encode(result)
}
) Prompts, completions, files, secrets, and inference stay local. There's no network path from inference, so Elix works with Wi-Fi off.
Only counts and durations leave the Mac over pinned TLS. Your prompts and completions don't.
Models are downloaded and SHA-256 verified. Once they're on the Mac, Elix is ready to work offline.
Every client is signed and validated before the service accepts its requests.
Elix runs Liquid Foundation Models through an MLX-based engine built for Apple silicon. It optimizes each model and tunes it for every chip we support.
How it's built
Built directly on MLX - unified memory, Metal kernels, and pure Swift. No Python runtime overhead between tokens.
Elix supports a focused set of models, so each one gets dedicated optimization - from custom Metal kernels to tuned quantization.
Every Apple silicon generation gets its own configuration, selected through measurements on real Macs.
What that gives you
Prefill tuning and model cache reuse reduce the wait before output begins.
Custom Metal kernels and speculative decoding increase generation throughput.
Inference runs on the Mac, with no server latency, rate limits, or cloud cold starts.
Decode and prefill throughput against the major Apple silicon engines - on real Macs, at prompt lengths from 128 to 32K tokens.
Prefill and decode combined: fastest in 19 runs, within 5% of the leader in 4 more.
Than the average of 6 rival engines, over all 25 chip-and-model runs and every prompt length.
Than the average of 6 rival engines at 32K-token prompts, across all 25 runs.
Decode and prefill measured, median tok/s at each prompt length, ranked by the geometric mean over 128–32K-token contexts. Time to first token is derived as prompt tokens over prefill throughput; end-to-end adds 256 generated tokens at decode throughput. Identical model weights across engines where the format allows: uzu runs its own M-format conversion, llama.cpp runs Q4_0. Each engine was benchmarked through its embedded OpenAI-compatible server, with guidellm as an independent profiling tool. Last updated in September 2026.
Same hardware, identical models, independent profiler. Only the engine changes from row to row.
Benchmarked on real Macs across different workloads, not extrapolated from a single run.
Speed comes from the engine, not from trading away model quality - outputs stay the model's own.
The model and the runtime work better when they're built with each other in mind. That's the idea behind our work with Liquid AI.
These aren't off-the-shelf checkpoints. We develop the models around the tasks Mac products need to run.
Liquid Foundation Models are efficient by architecture, then quantized and tuned for Elix's kernels on Apple silicon.
Models, inference, and memory are developed as one stack. Eney is the first MacPaw product running on it.
Some jobs just need a fast, focused model. Elix runs compact Liquid AI encoders for tasks like detecting sensitive data, routing prompts, choosing tools, and checking policies. All as steps in the same pipeline.
Spots 40 kinds of personal data across 16 languages and redacts them in one bidirectional pass - no generation loop, nothing leaves the Mac.
Reschedule my 3pm call with John Smith and text him at +1 234 567 890. Then pay the studio invoice with card 4111 1111 1111 1111 and email the receipt to [email protected] from [email protected].
Elix brings pipelines, agents, shared inference, and on-device models together for common Mac AI use cases.
Eney is where the Elix stack comes together: reasoning runs as pipelines through the shared service, with Mnemos providing memory - all on the user's Mac.
Build loops, tool calls, and branching with the same pipeline constructs as any other workflow. The open-source repository includes a working coding agent built this way.
Apps share one service and one resident model. Requests are batched across clients, and when all apps go idle, the model unloads completely.
Build summarization, structured extraction, and classification as declarative pipelines running on shared Liquid Foundation Models - locally and without per-query cloud costs.
The short answers - the rest of the page has the detail.
Elix is MacPaw's on-device AI runtime for the Mac, currently in development. You describe what your app needs to do in Swift, and Elix handles the rest - composable pipelines, one shared service across apps, and MLX-based inference tuned for Liquid Foundation Models on Apple silicon.
The pipeline layer is. Composable Pipelines - Elix's authoring, IR, and compiler layer - is on GitHub under Apache 2.0, with an OpenAI-compatible executor so you can run pipelines with Ollama or mlx-lm today. The Elix runtime itself is closed source.
Inference runs on the Mac - there is no network path from it, so Elix works with Wi-Fi off. Telemetry carries counts and durations, never prompts or completions. Verified model downloads are the one network hop in the system.
Elix is in development, and Eney is the first MacPaw product running on the stack. Join the waitlist for early access, or write to [email protected] to talk about a partnership.
Building something that could use on-device AI? Tell us what you're working on. Join the waitlist for early access to Elix, or get in touch to explore a partnership.
Join the waitlist ↗