Home / News / Perplexity Open‑Sources Lily: Run 35B LLM Locally on Apple Silicon
Perplexity Open‑Sources Lily: Run 35B LLM Locally on Apple Silicon
Perplexity AI has open‑sourced Lily, a Rust‑based inference engine that lets the 35‑billion‑parameter Qwen3.6‑35B‑A3B model run locally on Apple Silicon. Small businesses can now host cutting‑edge AI on their own Macs, cutting cloud costs, boosting privacy, and enhancing local search visibility.
Key Highlights
- ✓Perplexity opens source Lily inference engine
- ✓Runs 35B Qwen3.6 model on Apple Silicon
- ✓Local inference cuts cloud costs and boosts privacy
- ✓Empowers small businesses to control data and brand voice
What Happened
Perplexity AI has just released Lily, an open‑source inference engine that lets the 35‑billion‑parameter Qwen3.6‑35B‑A3B model run directly on Apple Silicon. This move removes the dependence on cloud providers and gives small businesses the ability to host the same cutting‑edge model on their own Macs.
Key Details
- Open‑source license – Lily ships under an MIT‑style license on GitHub, giving developers full freedom to modify, extend, or bundle the engine with their own applications.
- Apple‑Silicon optimized – Written in Rust for memory safety and accelerated with Metal for GPU usage, Lily delivers near‑native performance on M1, M2, and the upcoming M3 chips.
- Model support – The engine is engineered around the Qwen3.6‑35B‑A3B model, but its modular architecture can be adapted to other large language models, including open‑source variants.
- Local inference – Running inference locally eliminates the need for continuous cloud connectivity, protects sensitive data, and removes per‑token inference costs.
- Scalable architecture – While optimized for a single machine, the codebase is designed to scale across multiple Macs or a small cluster if needed.
What It Means For Your Business
1. Full Control Over Data and Brand Voice
Hosting the model on your own hardware keeps every user interaction in‑house. This is essential for compliance with privacy regulations such as GDPR or CCPA and lets you fine‑tune the model on proprietary data to match your brand’s tone and local knowledge.
2. Significant Cost Savings
Cloud inference can run into thousands of dollars per month for high‑volume queries. With Lily, you pay only for the Apple Silicon Mac you already own or plan to buy. After the initial hardware investment, the marginal cost of each query is essentially zero.
3. Lower Latency, Better User Experience
Because the model runs on the same device that serves the user, response times drop from several seconds (cloud) to less than a second. Faster answers translate into higher engagement, more conversions, and a better overall reputation in local search results.
4. Competitive Differentiation in AI‑Driven Search
Local businesses that can showcase AI‑powered features—such as instant FAQs, personalized recommendations, or instant content generation—stand out in search engine result pages (SERPs) and in the emerging AI‑centric overviews that power tools like ChatGPT and Gemini. Running the engine locally demonstrates technical sophistication and a commitment to privacy.
How To Get Started
1. Check your hardware – Ensure you have an Apple Silicon Mac with at least 16 GB of RAM and a recent macOS version that supports Metal.
2. Clone the repository – git clone https://github.com/perplexity-ai/lily.
3. Build the engine – Follow the Rust + Metal build instructions in the README. The process takes a few minutes on a modern Mac.
4. Download the Qwen3.6‑35B‑A3B model weights – These are available from the Qwen project under an open‑source license.
5. Run a test query – Use the provided CLI to send a prompt and see the response.
6. Integrate – Wrap the engine in a lightweight API layer and hook it into your local search widget, chatbot, or content‑generation pipeline.
7. Fine‑tune – If you have proprietary data (e.g., FAQs, product descriptions), fine‑tune the model to improve relevance.
Future Outlook
Perplexity’s open‑source move is part of a larger trend toward edge AI for search. As more companies release inference engines that run on consumer hardware, local businesses will gain unprecedented access to the same AI capabilities that large enterprises enjoy. This democratization of AI will likely accelerate the adoption of AI‑powered search overviews, making it essential for local brands to stay ahead of the curve.
By embracing Lily, small businesses can not only reduce costs and improve privacy but also position themselves as leaders in the next wave of AI‑driven discovery. The ability to host a 35‑billion‑parameter model on a single Mac is a game‑changer that will reshape how local search, citations, and AI assistants interact with your brand.
Why It Matters
Running a 35‑billion‑parameter LLM locally means your business can provide instant, AI‑driven answers without exposing customer data to third‑party cloud services. For local brands, this translates to faster response times, higher trust, and compliance with privacy regulations that increasingly influence search engine rankings.
Moreover, local inference eliminates the recurring costs that can drain marketing budgets. By investing in a single Apple Silicon machine, you gain a scalable AI engine that can power chatbots, content creation, and search optimization tools—all while keeping the data—and the insights—within your own network. As AI becomes a core component of how search engines surface local businesses, having your own inference engine positions you ahead of competitors who rely solely on cloud APIs.
FAQs
- What is Lily? Lily is an open‑source inference engine written in Rust and accelerated with Metal, designed to run large language models like Qwen3.6‑35B‑A3B on Apple Silicon.
- Do I need a Mac with an M1/M2 chip? Yes, Lily is optimized for Apple Silicon. A recent Mac with at least 16 GB of RAM is recommended for best performance.
- Will this help my local business appear in AI‑powered search results? Yes. By running a powerful LLM locally, you can power AI chatbots, instant FAQs, and content generation that search engines like ChatGPT and Gemini use to surface relevant local listings.
Why This Matters For Your Business
Running a 35‑billion‑parameter LLM locally means your business can deliver instant, AI‑driven answers without sending customer data to third‑party cloud services. For local brands, this translates into faster response times, higher trust, and compliance with privacy regulations that increasingly influence search engine rankings. Local inference also removes the recurring costs that can drain marketing budgets. A single Apple Silicon machine becomes a scalable AI engine capable of powering chatbots, content creation, and search‑optimization tools—all while keeping data and insights within your own network. As AI becomes a core component of how search engines surface local businesses, owning your own inference engine positions you ahead of competitors who rely solely on cloud APIs.
Frequently Asked Questions
What is Lily?
Lily is an open‑source inference engine written in Rust and accelerated with Metal, designed to run large language models like Qwen3.6‑35B‑A3B on Apple Silicon.
Do I need a Mac with an M1/M2 chip?
Yes, Lily is optimized for Apple Silicon. A recent Mac with at least 16 GB of RAM is recommended for best performance.
Will this help my local business appear in AI‑powered search results?
Yes. By running a powerful LLM locally, you can power AI chatbots, instant FAQs, and content generation that search engines like ChatGPT and Gemini use to surface relevant local listings.
Is your business showing up in AI search?
Get your free AI visibility audit - see if ChatGPT, Perplexity, and Google AI actually recommend you.
