✨ From vibe coding to vibe deployment. UBOS MCP turns ideas into infra with one message.

Learn more
Andrii Bidochko
  • Updated: March 20, 2026
  • 5 min read

HomeSec‑Bench Edge AI Benchmark News

HomeSec‑Bench proves that cutting‑edge LLMs can run on edge devices with performance rivaling cloud giants, delivering sub‑second time‑to‑first‑token and high decode speeds while keeping data fully private.

The benchmark, released by SharpAI, evaluates a suite of 96 tests across 15 real‑world security‑assistant scenarios, and the results show that a 9‑billion‑parameter model on a MacBook Pro M5 can achieve a 93.8 % pass rate—just four points shy of the leading cloud model.

Why Edge AI Benchmarks Matter Now

Edge computing is no longer a niche for hobbyists; it’s becoming the backbone of privacy‑first AI deployments in homes, factories, and remote sites. The SharpAI benchmark provides the most comprehensive, domain‑specific evaluation for LLMs on constrained hardware, focusing on metrics that matter to developers: GPU memory usage, time‑to‑first‑token (TTFT), and decode speed (tokens per second).

For tech enthusiasts, AI researchers, and decision‑makers, understanding these numbers is essential for choosing the right model‑hardware combo that balances cost, latency, and data sovereignty.

HomeSec‑Bench: An Overview

HomeSec‑Bench is a purpose‑built benchmark that simulates a home‑security AI assistant. Unlike generic chat tests, it measures how well a model can:

  • Classify security events (e.g., “masked person at night”).
  • Deduplicate alerts across multiple cameras.
  • Select and invoke the correct tool with proper parameters.
  • Resist prompt injection and maintain privacy compliance.

The suite comprises 96 LLM tests and 35 VLM (Vision‑Language Model) tests, organized into 15 functional categories ranging from Tool Use to Privacy & Compliance. All tests run against any OpenAI‑compatible endpoint, making the benchmark agnostic to the underlying inference engine.

Key Hardware Specs & Test Categories

Hardware Platform

CPU Apple M5 Pro (18 cores)
Unified Memory 64 GB
OS macOS 15.3 (arm64)
Inference Engine llama.cpp

Test Suites (selected)

  • Context Preprocessing – 6 tests
  • Topic Classification – 4 tests
  • Event Deduplication – 8 tests
  • Tool Use – 16 tests
  • Security Classification – 12 tests
  • Prompt Injection Resistance – 4 tests
  • Multi‑Turn Reasoning – 4 tests
  • Privacy & Compliance – 3 tests

Leaderboard Summary & Core Metrics

The leaderboard ranks models by overall pass rate, then by execution time. Below is a distilled view of the top performers.

Rank Model Deployment Pass Rate TTFT (ms) Decode Speed (tok/s) GPU Memory (GB)
🥇 GPT‑5.4 (cloud) ☁️ Cloud 97.9 % 2 200 234.5
🥈 GPT‑5.4‑mini (cloud) ☁️ Cloud 95.8 % 1 800 136.4
🥉 Qwen3.5‑9B (Q4_K_M) 🏠 Local 93.8 % 765 25 13.8
🏅 Qwen3.5‑35B‑MoE (Q4_K_L) 🏠 Local 91.7 % 435 41.9 27.2

Key takeaways:

  • The 9B Qwen model runs entirely offline on a single laptop, using only 13.8 GB of unified memory.
  • Its TTFT of 765 ms is competitive with many cloud offerings, while decode speed of 25 tok/s is sufficient for real‑time alert triage.
  • Higher‑parameter MoE models (35B) achieve sub‑500 ms TTFT but require significantly more memory, highlighting the classic trade‑off between latency and resource consumption.

Insights for Edge AI Practitioners

HomeSec‑Bench delivers more than raw numbers; it surfaces strategic insights for anyone building AI on the edge.

1. Memory Footprint Drives Feasibility

Models that stay under 15 GB of GPU memory can comfortably run on most modern laptops and edge servers. The Qwen3.5‑9B’s 13.8 GB usage makes it a sweet spot for developers who need privacy without provisioning expensive GPUs.

2. TTFT Is the Real‑World Bottleneck

In security‑alert pipelines, the first token often determines whether a user sees an alert in time. Sub‑500 ms TTFT, as demonstrated by the Qwen3.5‑35B‑MoE, can enable “instant” response scenarios, while still keeping the model on‑device.

3. Decode Speed Aligns With Task Complexity

Simple classification tasks (e.g., “Is this person masked?”) need only a few tokens, so 25 tok/s is ample. More complex tool‑selection or multi‑turn reasoning benefits from higher decode speeds; here, the 41.9 tok/s of the 35B‑MoE shines.

4. Cost‑Effective Privacy

Running locally eliminates API fees and data‑exfiltration risks. For enterprises with strict compliance mandates, the ability to achieve a 93.8 % pass rate without any cloud dependency is a game‑changer.

5. Integration Pathways with UBOS

Developers can accelerate deployment by leveraging the UBOS platform overview, which offers pre‑built connectors for LLM inference, edge orchestration, and monitoring. The platform’s edge‑computing capabilities simplify scaling from a single laptop to a fleet of edge nodes.

HomeSec‑Bench performance chart
Figure 1 – Comparative performance of top LLMs on the HomeSec‑Bench benchmark (TTFT, decode speed, and memory usage).

Conclusion: Edge AI Is Ready for Production

HomeSec‑Bench demonstrates that high‑quality LLM inference is no longer the exclusive domain of massive cloud clusters. With a 9B model delivering near‑cloud accuracy on a consumer‑grade MacBook, organizations can now deploy privacy‑first AI assistants at the edge, cut operational costs, and meet stringent compliance requirements.

Ready to experiment with your own edge AI workflows? Explore the UBOS homepage for a free trial, or dive into the edge‑computing suite to provision models on your hardware in minutes.

For marketers looking to amplify AI‑driven campaigns, the AI marketing agents can automatically generate copy, analyze performance, and adapt messaging based on real‑time insights—leveraging the same LLM technology proven in HomeSec‑Bench.

Stay ahead of the curve: adopt edge AI today, protect your data tomorrow.


Andrii Bidochko

CTO UBOS

Andrii Bidochko is an AI entrepreneur and researcher focused on AI agents, reinforcement learning, and autonomous systems. He writes about the technologies shaping the future of machine intelligence, from frontier models and agent architectures to real-world AI applications.

Sign up for our newsletter

Stay up to date with the roadmap progress, announcements and exclusive discounts feel free to sign up with your email.

Sign In

Register

Reset Password

Please enter your username or email address, you will receive a link to create a new password via email.