S Syvidea

All options are based on the same AMD Strix Halo platform. Differences lie in workload optimization, not raw performance.

Architecture

AMD Strix Halo architecture

A system-on-chip design built around a large unified memory pool, where the CPU, GPU, and NPU access the same high-bandwidth memory. This is the technical foundation of the workstations Syvidea sources.

This architecture is the reason Syvidea can unify three different workstation brands under one comparison system.

Architecture value

Why this design matters for your AI workload

Unified memory design

Designed for efficient large model inference without splitting memory between CPU and GPU.

High bandwidth

Optimized for memory-bound LLM and RAG workloads, reducing token latency.

Single platform

Mapped to three partner brands with consistent performance baseline.

Core idea

Unified memory changes the local AI trade-off

Traditional desktops separate system RAM from GPU VRAM. A discrete GPU may have 24 GB of VRAM while the system has 64 GB of RAM, and data must be copied across the PCIe bus. Strix Halo places up to 128 GB of LPDDR5X memory on a wide 256-bit interface, shared by the CPU, the Radeon 8060S integrated GPU, and the XDNA 2 NPU.

For local AI, this means a large language model can reside in one memory pool without being split between "CPU memory" and "GPU memory." The bottleneck becomes memory bandwidth, not PCIe transfer or VRAM capacity.

What this means for you:

You can run larger models locally than what's typically possible with discrete GPU workstations at this price point. The unified memory eliminates the need to worry about VRAM capacity limits.

Compute units

CPU, GPU, and NPU interaction

CPU

Zen 5 cores

16 cores / 32 threads provide general-purpose compute for preprocessing, orchestration, and workloads that are not heavily parallelized.

iGPU

Radeon 8060S

RDNA 3.5 integrated graphics handles matrix-heavy inference, image generation, and multimodal tasks with direct access to the full unified memory pool.

NPU

XDNA 2

Up to 50 TOPS of INT8 throughput for supported inference, background AI tasks, and power-efficient model execution.

Memory bandwidth

Bandwidth vs. compute trade-off

Strix Halo is a memory-bandwidth-bound platform for large-model inference. A 70B-parameter model loaded in 4-bit quantization needs roughly 35–40 GB of memory just for weights, plus additional space for context, KV cache, and overhead. The 128 GB unified memory pool leaves room for these workloads.

However, the raw compute throughput is lower than a desktop NVIDIA RTX 4090 or 5090 GPU. Strix Halo is optimized for memory capacity and unified access, not for maximum tokens per second or CUDA-specific image-generation pipelines.

What this means for you:

You should prioritize this platform for memory-heavy workloads like RAG systems and large LLM inference, not for maximum-speed CUDA tasks like Flux image generation.