All options are based on the same AMD Strix Halo platform. Differences lie in workload optimization, not raw performance.
AMD Strix Halo architecture
A system-on-chip design built around a large unified memory pool, where the CPU, GPU, and NPU access the same high-bandwidth memory. This is the technical foundation of the workstations Syvidea sources.
This architecture is the reason Syvidea can unify three different workstation brands under one comparison system.
Why this design matters for your AI workload
Designed for efficient large model inference without splitting memory between CPU and GPU.
Optimized for memory-bound LLM and RAG workloads, reducing token latency.
Mapped to three partner brands with consistent performance baseline.
Unified memory changes the local AI trade-off
Traditional desktops separate system RAM from GPU VRAM. A discrete GPU may have 24 GB of VRAM while the system has 64 GB of RAM, and data must be copied across the PCIe bus. Strix Halo places up to 128 GB of LPDDR5X memory on a wide 256-bit interface, shared by the CPU, the Radeon 8060S integrated GPU, and the XDNA 2 NPU.
For local AI, this means a large language model can reside in one memory pool without being split between "CPU memory" and "GPU memory." The bottleneck becomes memory bandwidth, not PCIe transfer or VRAM capacity.
What this means for you:
You can run larger models locally than what's typically possible with discrete GPU workstations at this price point. The unified memory eliminates the need to worry about VRAM capacity limits.
CPU, GPU, and NPU interaction
Zen 5 cores
16 cores / 32 threads provide general-purpose compute for preprocessing, orchestration, and workloads that are not heavily parallelized.
Radeon 8060S
RDNA 3.5 integrated graphics handles matrix-heavy inference, image generation, and multimodal tasks with direct access to the full unified memory pool.
XDNA 2
Up to 50 TOPS of INT8 throughput for supported inference, background AI tasks, and power-efficient model execution.
Bandwidth vs. compute trade-off
Strix Halo is a memory-bandwidth-bound platform for large-model inference. A 70B-parameter model loaded in 4-bit quantization needs roughly 35–40 GB of memory just for weights, plus additional space for context, KV cache, and overhead. The 128 GB unified memory pool leaves room for these workloads.
However, the raw compute throughput is lower than a desktop NVIDIA RTX 4090 or 5090 GPU. Strix Halo is optimized for memory capacity and unified access, not for maximum tokens per second or CUDA-specific image-generation pipelines.
What this means for you:
You should prioritize this platform for memory-heavy workloads like RAG systems and large LLM inference, not for maximum-speed CUDA tasks like Flux image generation.