{"id":1573,"date":"2026-09-28T23:00:15","date_gmt":"2026-09-28T23:00:15","guid":{"rendered":"https:\/\/yairmartinezcybersecurityportfolio.com\/?p=1573"},"modified":"2026-09-29T06:16:17","modified_gmt":"2026-09-29T06:16:17","slug":"local-ai-foundation-building-a-self-hosted-multi-gpu-ai-homelab","status":"publish","type":"post","link":"https:\/\/yairmartinezcybersecurityportfolio.com\/?p=1573","title":{"rendered":"Local AI Foundation \u2014 Building a Self-Hosted Multi-GPU AI Homelab"},"content":{"rendered":"\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">Introduction<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">This project documents the first major foundation of my local AI homelab. The goal was not just to install a chatbot or download a single model. I wanted to build a self-hosted AI environment that works more like real infrastructure: separated services, documented storage, recoverable backups, and different backends for different use cases.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The current system uses Proxmox as the host platform and runs three main Linux containers:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>CT100:<\/strong> OpenWebUI and Ollama for the main user-facing AI interface<\/li>\n\n\n\n<li><strong>CT101:<\/strong> llama.cpp for a stable OpenAI-compatible GGUF backend<\/li>\n\n\n\n<li><strong>CT102:<\/strong> vLLM as a proof-of-concept for serving Qwen3.8-27B-FP8 across multiple GPUs<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">I built this in parts so I could understand the system instead of treating it as a black box. Each container has a specific role, and the storage\/backups are designed so the system can be rebuilt later if something fails.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The larger goal is to turn this machine into a private multimodal AI workstation over time. For now, this post focuses on the infrastructure foundation: the host, the containers, the models, the performance notes, the troubleshooting, and the backup strategy.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">Project Goals<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">My main goals for this project were:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Build a local AI system on my own hardware<\/li>\n\n\n\n<li>Separate the main AI services into containers<\/li>\n\n\n\n<li>Use consumer NVIDIA GPUs for local inference experiments<\/li>\n\n\n\n<li>Test practical backends like Ollama and llama.cpp<\/li>\n\n\n\n<li>Experiment with vLLM, FP8, and multi-GPU serving<\/li>\n\n\n\n<li>Keep large model files outside the container root disks<\/li>\n\n\n\n<li>Create hot\/warm backups on local storage<\/li>\n\n\n\n<li>Plan for cold storage backups on an external drive<\/li>\n\n\n\n<li>Document the setup well enough that I could rebuild it later<\/li>\n\n\n\n<li>Turn the build into a portfolio project instead of leaving it as undocumented experimentation<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">I also wanted the system to support future projects, especially a Hermes-style local agent workflow, cybersecurity\/IT support assistants, and eventually local multimodal workflows such as image, speech, transcription, and video processing.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">Visual Stack Model<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">This is the current stack at a high level:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>User \/ Browser \/ Main PC\n        |\n        v\n+------------------------------+\n| CT100: OpenWebUI + Ollama    |\n| - Main web interface         |\n| - Daily-use AI frontend      |\n| - Ollama model backend       |\n+--------------+---------------+\n               |\n               | OpenAI-compatible \/ backend connections\n               v\n+------------------------------+       +------------------------------+\n| CT101: llama.cpp             |       | CT102: vLLM                  |\n| - Qwen3.8 GGUF backend       |       | - Qwen3.8 FP8 proof of       |\n| - Stable API backend         |       |   concept                    |\n| - Agent-backend candidate    |       | - Short-context testing      |\n+------------------------------+       +------------------------------+\n\nHost Platform:\n+--------------------------------------------------------------+\n| prox0 - Proxmox VE                                           |\n| - Ryzen 9 5900X                                              |\n| - RTX 5060 Ti 16GB + RTX 3060 12GB + RTX 3060 12GB           |\n| - NVMe ZFS mirror for Proxmox and CT root disks              |\n| - SATA AI storage for models, cache, docs, and warm backups  |\n+--------------------------------------------------------------+\n\nStorage and Recovery:\n+---------------------+     +----------------------------+\n| NVMe ZFS mirror     |     | \/mnt\/ai-sata SATA storage  |\n| - Proxmox OS        |     | - Ollama models            |\n| - CT root disks     |     | - GGUF models              |\n|                     |     | - vLLM HF cache            |\n|                     |     | - warm CT backups          |\n|                     |     | - documentation            |\n+---------------------+     +----------------------------+\n                                      |\n                                      v\n                           External HDD \/ Cold Storage\n                           - disaster recovery copy<\/code><\/pre>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"531\" src=\"https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-205804-1024x531.png\" alt=\"\" class=\"wp-image-1582\" srcset=\"https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-205804-1024x531.png 1024w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-205804-300x156.png 300w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-205804-767x398.png 767w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-205804-439x228.png 439w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-205804-678x352.png 678w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-205804-1536x797.png 1536w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-205804.png 1607w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"306\" height=\"236\" src=\"https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-205932.png\" alt=\"\" class=\"wp-image-1583\" srcset=\"https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-205932.png 306w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-205932-300x231.png 300w\" sizes=\"auto, (max-width: 306px) 100vw, 306px\" \/><\/figure>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">Part 1 \u2014 Host Hardware and Proxmox Foundation<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The host is named <strong>prox0<\/strong> and runs Proxmox VE. It is built on a Ryzen 9 5900X with about 32 GB of RAM and three NVIDIA GPUs:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>RTX 5060 Ti 16 GB<\/li>\n\n\n\n<li>RTX 3060 12 GB<\/li>\n\n\n\n<li>RTX 3060 12 GB<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">That gives the rig about 40 GB of installed VRAM across the system. One important point I learned is that total VRAM is not the same thing as one large shared memory pool. Each backend handles multi-GPU workloads differently, and some models split better than others.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The current philosophy is to run <strong>one heavy AI workload at a time<\/strong>. This is not a production cluster meant to keep every model loaded at once. It is a local AI workstation where I can choose the best backend for the job.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"545\" src=\"https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-210125-1024x545.png\" alt=\"\" class=\"wp-image-1586\" srcset=\"https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-210125-1024x545.png 1024w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-210125-300x160.png 300w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-210125-767x408.png 767w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-210125-440x234.png 440w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-210125-678x361.png 678w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-210125.png 1383w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">Part 2 \u2014 Storage Design<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">I separated the storage into two main layers:<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">NVMe ZFS Mirror<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">The NVMe mirror is used for:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Proxmox OS<\/li>\n\n\n\n<li>Container root disks<\/li>\n\n\n\n<li>Core system state<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">This gives the host and container root filesystems better protection than a single boot drive.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">SATA AI Storage<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Large AI files live on a SATA SSD mounted at:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>\/mnt\/ai-sata<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">This storage holds:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Ollama model store<\/li>\n\n\n\n<li>GGUF models<\/li>\n\n\n\n<li>vLLM Hugging Face cache<\/li>\n\n\n\n<li>vLLM cache<\/li>\n\n\n\n<li>Proxmox warm backups<\/li>\n\n\n\n<li>Documentation files<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Important folders:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>\/mnt\/ai-sata\/ollama\n\/mnt\/ai-sata\/gguf\n\/mnt\/ai-sata\/vllm\n\/mnt\/ai-sata\/containers\/dump\n\/mnt\/ai-sata\/docs<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">This design keeps large model files out of the CT root disks. It also makes the backup plan more obvious: the container backup protects the container configuration and service environment, but the model folders need to be copied separately.<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"755\" height=\"376\" src=\"https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-210533.png\" alt=\"\" class=\"wp-image-1587\" srcset=\"https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-210533.png 755w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-210533-300x149.png 300w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-210533-440x219.png 440w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-210533-679x338.png 679w\" sizes=\"auto, (max-width: 755px) 100vw, 755px\" \/><\/figure>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">Part 3 \u2014 CT100: OpenWebUI and Ollama<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">CT100 is the main user-facing AI container. It runs OpenWebUI and Ollama.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">OpenWebUI gives me the browser-based interface, while Ollama handles local model serving. This is the easiest part of the stack to use every day.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Main role:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Main local AI interface<\/li>\n\n\n\n<li>Ollama backend<\/li>\n\n\n\n<li>Daily-use candidate<\/li>\n\n\n\n<li>Model storage through a SATA bind mount<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">The Ollama model store is mounted from the host into the container:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>Host path:      \/mnt\/ai-sata\/ollama\nContainer path: \/mnt\/ollama-models<\/code><\/pre>\n\n\n\n<h3 class=\"wp-block-heading\">Installed Ollama Models<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">At the time of documentation, the installed Ollama models were:<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Model<\/th><th>Size<\/th><th>Intended Role<\/th><\/tr><\/thead><tbody><tr><td><code>qwen3.8:27b<\/code><\/td><td>17 GB<\/td><td>Main general-purpose local model for chat, reasoning, coding, research, and documentation help<\/td><\/tr><tr><td><code>muse-glimmer:30b<\/code><\/td><td>18 GB<\/td><td>Creative\/image-related model candidate and future visual workflow testing<\/td><\/tr><tr><td><code>nemotron-3.5-lightning:30b-a3b-q4_K_M<\/code><\/td><td>25 GB<\/td><td>Experimental alternate model, not necessarily a long-term keeper<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">The main daily-driver model direction is <strong>Qwen3.8 27B<\/strong>. My goal is not to collect many similar models. I want each model to earn a role.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Observed Performance<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">In testing, Ollama with Qwen3.8 27B was the strongest daily-use result. A benchmark run showed roughly:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>Ollama Qwen3.8 27B\nPrompt eval: ~1,862 tokens\/sec\nGeneration:  ~18.5 tokens\/sec<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">That made CT100 the easiest and most practical day-to-day interface.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"485\" src=\"https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-220448-1024x485.png\" alt=\"\" class=\"wp-image-1589\" srcset=\"https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-220448-1024x485.png 1024w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-220448-300x142.png 300w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-220448-767x363.png 767w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-220448-1536x727.png 1536w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-220448-439x208.png 439w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-220448-678x321.png 678w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-220448.png 1916w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"316\" height=\"174\" src=\"https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/image.png\" alt=\"\" class=\"wp-image-1591\" srcset=\"https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/image.png 316w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/image-300x165.png 300w\" sizes=\"auto, (max-width: 316px) 100vw, 316px\" \/><\/figure>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">Part 4 \u2014 CT101: llama.cpp Backend<\/h2>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1280\" height=\"613\" src=\"https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/ezgif.com-optimize-1.gif\" alt=\"\" class=\"wp-image-1597\"\/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">CT101 runs llama.cpp and serves a Qwen3.8 GGUF model through <code>llama-server<\/code>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This container matters because llama.cpp gives me a stable OpenAI-compatible API endpoint. That makes it useful for OpenWebUI, future tools, and possible agent frameworks.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Main role:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Stable OpenAI-compatible backend<\/li>\n\n\n\n<li>GGUF model serving<\/li>\n\n\n\n<li>Alternate daily-driver backend<\/li>\n\n\n\n<li>Strong candidate for future Hermes\/agent integration<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Main model:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>\/models\/qwen3.8-27b\/Qwen3.8-27B-Q4_K_M.gguf<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">The model is stored on the SATA-backed <code>\/models<\/code> bind mount:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>Host path:      \/mnt\/ai-sata\/gguf\nContainer path: \/models<\/code><\/pre>\n\n\n\n<h3 class=\"wp-block-heading\">llama.cpp Performance Notes<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">The llama.cpp backend performed well enough to be considered a real daily-driver candidate. In testing, the Qwen3.8 GGUF backend produced results around:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>llama.cpp Qwen3.8 GGUF\nPrompt eval: ~717 tokens\/sec\nGeneration:  ~15.6 tokens\/sec<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Other runs showed generation in the mid-to-high 16 tokens\/sec range depending on the prompt and run conditions. The bigger takeaway is that CT101 was stable and practical, especially compared with the vLLM proof-of-concept.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Notable llama.cpp Troubleshooting<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">One issue I documented was a <code>503 Loading model<\/code> API response immediately after service startup. That was not a failure. It meant the llama.cpp server had started, but the model was still loading. After startup finished, the API could be tested normally through <code>\/v1\/models<\/code> and <code>\/v1\/chat\/completions<\/code>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This was a useful reminder that service status and model readiness are not always the same thing. A systemd service can be running while the model is still warming up.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"537\" src=\"https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-234512-1024x537.png\" alt=\"\" class=\"wp-image-1599\" srcset=\"https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-234512-1024x537.png 1024w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-234512-300x157.png 300w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-234512-439x230.png 439w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-234512-767x402.png 767w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-234512-679x356.png 679w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-234512.png 1106w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"539\" src=\"https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-225558-1024x539.png\" alt=\"\" class=\"wp-image-1600\" srcset=\"https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-225558-1024x539.png 1024w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-225558-300x158.png 300w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-225558-768x404.png 768w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-225558-439x231.png 439w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-225558-678x357.png 678w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-225558.png 1393w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"307\" height=\"211\" src=\"https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-224017.png\" alt=\"\" class=\"wp-image-1593\" srcset=\"https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-224017.png 307w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-28-224017-300x206.png 300w\" sizes=\"auto, (max-width: 307px) 100vw, 307px\" \/><\/figure>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">Part 5 \u2014 CT102: vLLM Proof of Concept<\/h2>\n\n\n\n<figure class=\"wp-block-image size-full is-resized\"><img loading=\"lazy\" decoding=\"async\" width=\"800\" height=\"381\" src=\"https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/ezgif.com-optimize-2.gif\" alt=\"\" class=\"wp-image-1603\" style=\"width:980px;height:auto\"\/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">CT102 is the vLLM experiment container. It is not my main daily-driver backend.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The goal of CT102 was to prove that the rig could serve <strong>Qwen3.8-27B-FP8<\/strong> through vLLM using all three NVIDIA GPUs with pipeline parallelism.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Main role:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>vLLM proof-of-concept<\/li>\n\n\n\n<li>Qwen3.8-27B-FP8 testing<\/li>\n\n\n\n<li>OpenAI-compatible API endpoint<\/li>\n\n\n\n<li>Multi-GPU pipeline-parallel experiment<\/li>\n\n\n\n<li>Short-prompt demo backend<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Final served model:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>qwen38-27b-fp8<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Hugging Face model:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>Qwen\/Qwen3.8-27B-FP8<\/code><\/pre>\n\n\n\n<h3 class=\"wp-block-heading\">Final Stable vLLM Profile<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">The final working vLLM profile used:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>pipeline_parallel_size = 3\ntensor_parallel_size   = 1\nlayer split            = 20 \/ 20 \/ 24\ncontext                = 384 tokens\nKV cache               = fixed 512 MiB\nKV cache dtype         = fp8\nmode                   = text-only\neager execution        = enabled<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">The important result is that CT102 works as a proof-of-concept. It proved that the rig could load and serve the FP8 model across all three GPUs.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">vLLM Performance Notes<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">One successful vLLM run showed approximately:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>vLLM Qwen3.8-27B-FP8\nPrompt throughput:     ~39 tokens\/sec\nGeneration throughput: ~7.8 tokens\/sec<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">That result was useful, but it did not make vLLM the best daily workflow. The stable context was intentionally tiny, and the hardware layout created tradeoffs.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Notable vLLM Troubleshooting<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">This was the most troubleshooting-heavy part of the project.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Some approaches that did not work:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Tensor Parallel 3:<\/strong> failed because model dimensions\/vocab were not divisible by 3.<\/li>\n\n\n\n<li><strong>Tensor Parallel 2:<\/strong> could not fit cleanly on the two 12 GB GPUs without CPU offload.<\/li>\n\n\n\n<li><strong>CPU offload:<\/strong> technically helped memory pressure, but used too much RAM and hurt performance.<\/li>\n\n\n\n<li><strong>Larger context sizes:<\/strong> caused CUDA OOM, Marlin FP8 runtime failures, or EngineDead errors.<\/li>\n\n\n\n<li><strong>21\/21\/22 layer split:<\/strong> shifted the failure point instead of solving it.<\/li>\n\n\n\n<li><strong>Mixed consumer GPU daily-driver use:<\/strong> FP8 vLLM was not a good daily-driver fit for the 12 GB + 12 GB + 16 GB GPU layout.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">The final stable profile was built by accepting the limitation and changing the goal: CT102 became a clean proof-of-concept instead of a daily-driver backend.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"538\" src=\"https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-29-015058-1024x538.png\" alt=\"\" class=\"wp-image-1604\" srcset=\"https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-29-015058-1024x538.png 1024w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-29-015058-300x158.png 300w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-29-015058-1536x807.png 1536w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-29-015058-767x403.png 767w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-29-015058-440x231.png 440w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-29-015058-679x357.png 679w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-29-015058-2048x1076.png 2048w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"336\" src=\"https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-29-014858-1024x336.png\" alt=\"\" class=\"wp-image-1605\" srcset=\"https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-29-014858-1024x336.png 1024w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-29-014858-300x99.png 300w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-29-014858-767x252.png 767w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-29-014858-1536x504.png 1536w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-29-014858-439x144.png 439w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-29-014858-679x223.png 679w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-29-014858-2048x672.png 2048w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"731\" src=\"https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-29-015838-1024x731.png\" alt=\"\" class=\"wp-image-1606\" srcset=\"https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-29-015838-1024x731.png 1024w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-29-015838-300x214.png 300w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-29-015838-768x548.png 768w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-29-015838-680x485.png 680w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-29-015838-439x313.png 439w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-29-015838-1536x1096.png 1536w, https:\/\/yairmartinezcybersecurityportfolio.com\/wp-content\/uploads\/2026\/09\/Screenshot-2026-09-29-015838.png 1585w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">Consumer Hardware Tradeoffs<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">This project also showed the difference between consumer hardware that can run a workload and server hardware that is ideal for it.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The system has three GPUs, but the PCIe layout is not perfect for multi-GPU serving. The two RTX 3060 cards were observed running at downgraded <strong>PCIe x1<\/strong> links, while the RTX 5060 Ti was observed at a downgraded <strong>x8<\/strong> link. That does not stop the system from working, but it matters for workloads that need GPUs to communicate or pass data through pipeline stages.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The practical impact was clearest with vLLM. Pipeline parallelism across all three GPUs worked, but the x1 links on the RTX 3060s made this a better proof-of-concept than a daily-driver setup.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This changed how I think about the backend roles:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Ollama:<\/strong> best convenience and daily-use experience so far<\/li>\n\n\n\n<li><strong>llama.cpp:<\/strong> stable OpenAI-compatible backend and strong agent candidate<\/li>\n\n\n\n<li><strong>vLLM:<\/strong> valuable experiment and portfolio proof, but not the main workflow on this hardware<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">That lesson is important because it is easy to look only at total VRAM and miss the rest of the system design.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">Benchmark Snapshot<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">These numbers are not formal lab benchmarks. They are practical observations from my own testing and logs.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Backend<\/th><th>Model<\/th><th>Approx. Prompt Throughput<\/th><th>Approx. Generation Throughput<\/th><th>Practical Result<\/th><\/tr><\/thead><tbody><tr><td>Ollama<\/td><td>Qwen3.8 27B<\/td><td>~1,862 tokens\/sec<\/td><td>~18.5 tokens\/sec<\/td><td>Best daily-use result so far<\/td><\/tr><tr><td>llama.cpp<\/td><td>Qwen3.8 27B GGUF<\/td><td>~717 tokens\/sec<\/td><td>~15.6 tokens\/sec<\/td><td>Stable OpenAI-compatible backend<\/td><\/tr><tr><td>vLLM<\/td><td>Qwen3.8-27B-FP8<\/td><td>~39 tokens\/sec<\/td><td>~7.8 tokens\/sec<\/td><td>Working proof-of-concept, not daily driver<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">The benchmark results helped me make a practical decision. vLLM was interesting technically, but Ollama and llama.cpp were better fits for normal use on this rig.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">Backup and Recovery Plan<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Backup planning became one of the most important parts of this project.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The Proxmox CT backups protect the container root filesystems. That includes the installed services, configuration, scripts, and system state inside the containers.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">However, the large model folders are bind mounts. That means normal CT backups do not fully include the model directories.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">My simple way of thinking about the backup design is:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>CT backup           = container brain \/ config \/ root filesystem\nModel folder backup = large model bodies and cache\nDocumentation       = rebuild knowledge<\/code><\/pre>\n\n\n\n<h3 class=\"wp-block-heading\">Warm Backup Layer<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">The SATA SSD acts as the warm backup\/staging layer. It stores:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>\/mnt\/ai-sata\/containers\/dump\n\/mnt\/ai-sata\/docs<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">The container backups are stored under:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>\/mnt\/ai-sata\/containers\/dump<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">This gives me a local recovery point if I break a container or need to roll back the AI services.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Cold Storage Layer<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">The cold storage plan is to copy the important backup set to an external hard drive.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Cold storage should include:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>\/mnt\/ai-sata\/containers\/dump\n\/mnt\/ai-sata\/ollama\n\/mnt\/ai-sata\/gguf\n\/mnt\/ai-sata\/vllm\n\/mnt\/ai-sata\/docs<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">This matters because a CT backup alone is not enough. If the SATA SSD failed and I only had the CT root backups, I would lose the downloaded model stores and caches.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">Documentation Pack<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">I created a final documentation pack so the project can be understood later instead of relying on memory.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The documentation pack includes:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><code>README.txt<\/code><\/li>\n\n\n\n<li><code>ai-rig-overview.txt<\/code><\/li>\n\n\n\n<li><code>proxmox-host-raid1-nvme.txt<\/code><\/li>\n\n\n\n<li><code>ct100-openwebui-ai.txt<\/code><\/li>\n\n\n\n<li><code>ct101-llamacpp.txt<\/code><\/li>\n\n\n\n<li><code>ct102-vllm.txt<\/code><\/li>\n\n\n\n<li><code>ai-rig-roadmap.txt<\/code><\/li>\n\n\n\n<li><code>ai-rig-ideas.txt<\/code><\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">This made the project much more structured. The files document the host, storage design, container roles, installed models, backup rules, troubleshooting notes, and future roadmap.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">What I Want To Use This For<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The immediate use case is local AI chat and technical assistance. The longer-term goal is to build a private multimodal AI workstation.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Current and future uses include:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>General local assistant work<\/li>\n\n\n\n<li>Coding help<\/li>\n\n\n\n<li>Research and summarization<\/li>\n\n\n\n<li>Homelab troubleshooting<\/li>\n\n\n\n<li>Documentation generation<\/li>\n\n\n\n<li>Cybersecurity and IT support workflows<\/li>\n\n\n\n<li>Splunk and log-analysis assistance<\/li>\n\n\n\n<li>Windows Event Log review<\/li>\n\n\n\n<li>Suricata alert explanation<\/li>\n\n\n\n<li>PowerShell and Linux command support<\/li>\n\n\n\n<li>Agent workflow testing with Hermes<\/li>\n\n\n\n<li>Image-generation and thumbnail workflows<\/li>\n\n\n\n<li>Speech-to-text and local voice assistant experiments<\/li>\n\n\n\n<li>Local content-production workflows<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">The important design rule is that I do not want to collect random models just because they exist. For each category, I want to test and keep the best-performing model or tool that my rig can realistically run.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">Key Lessons Learned<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The biggest lessons from this project were practical, not theoretical.<\/p>\n\n\n\n<ol start=\"1\" class=\"wp-block-list\">\n<li><strong>Stable daily use matters more than theoretical performance.<\/strong><br>vLLM was the most complex and interesting backend, but Ollama and llama.cpp are more useful for daily work right now.<\/li>\n\n\n\n<li><strong>Total VRAM is not the whole story.<\/strong><br>GPU memory, PCIe link speed, backend architecture, model format, and quantization all affect the real result.<\/li>\n\n\n\n<li><strong>Consumer hardware can do serious work, but it has tradeoffs.<\/strong><br>The x1 links on the RTX 3060s made some multi-GPU workloads less practical, especially vLLM pipeline parallelism.<\/li>\n\n\n\n<li><strong>Bind mounts are useful but affect backups.<\/strong><br>Keeping models outside the CT root disks is cleaner, but those model folders need their own backup plan.<\/li>\n\n\n\n<li><strong>A service can be running before a model is ready.<\/strong><br>The llama.cpp <code>503 Loading model<\/code> response was a good example of why startup state and model readiness need to be checked separately.<\/li>\n\n\n\n<li><strong>Proof-of-concept and daily-driver are different goals.<\/strong><br>CT102 is successful because it proved vLLM could run the model. That does not mean it should be the main workflow.<\/li>\n\n\n\n<li><strong>Documentation turns experiments into infrastructure.<\/strong><br>Writing the documentation pack made the setup easier to explain, recover, and turn into a portfolio project.<\/li>\n<\/ol>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">Final Thoughts<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">This project helped me move from experimenting with local AI models to building an actual AI infrastructure foundation.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The most important result is not just that the models run. The important result is that the system now has a structure:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Proxmox host documented<\/li>\n\n\n\n<li>Storage layout documented<\/li>\n\n\n\n<li>AI containers separated by role<\/li>\n\n\n\n<li>Backends tested and compared<\/li>\n\n\n\n<li>Models stored outside CT root disks<\/li>\n\n\n\n<li>Warm backups created<\/li>\n\n\n\n<li>Cold storage plan defined<\/li>\n\n\n\n<li>Troubleshooting notes preserved<\/li>\n\n\n\n<li>Future roadmap written<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">For my current hardware, the best practical path is to use OpenWebUI\/Ollama and llama.cpp as the main daily-driver options, while keeping vLLM as a working proof-of-concept and learning environment.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This gives me a strong foundation for the next phase: building a local agent workflow and turning the rig into a more capable private multimodal AI workstation.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Introduction This project documents the first major foundation of my local AI homelab. The goal was not just to install a chatbot or download a single model. I wanted to build a self-hosted AI environment that works more like real infrastructure: separated services, documented storage, recoverable backups, and different backends for different use cases. The [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":1578,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[],"class_list":["post-1573","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-projects"],"_links":{"self":[{"href":"https:\/\/yairmartinezcybersecurityportfolio.com\/index.php?rest_route=\/wp\/v2\/posts\/1573","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/yairmartinezcybersecurityportfolio.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/yairmartinezcybersecurityportfolio.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/yairmartinezcybersecurityportfolio.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/yairmartinezcybersecurityportfolio.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=1573"}],"version-history":[{"count":9,"href":"https:\/\/yairmartinezcybersecurityportfolio.com\/index.php?rest_route=\/wp\/v2\/posts\/1573\/revisions"}],"predecessor-version":[{"id":1607,"href":"https:\/\/yairmartinezcybersecurityportfolio.com\/index.php?rest_route=\/wp\/v2\/posts\/1573\/revisions\/1607"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/yairmartinezcybersecurityportfolio.com\/index.php?rest_route=\/wp\/v2\/media\/1578"}],"wp:attachment":[{"href":"https:\/\/yairmartinezcybersecurityportfolio.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=1573"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/yairmartinezcybersecurityportfolio.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=1573"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/yairmartinezcybersecurityportfolio.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=1573"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}