Skip to content

LambertLab Infrastructure Documentation

Welcome to the central technical documentation and operational engineering knowledge base for the LambertLab hybrid cloud-native infrastructure.


⚡ Architecture Topology

Hold "Alt" / "Option" to enable pan & zoom
graph TD
    User["User / External Client"] --> CF["Cloudflare Multi-A DNS<br/>*.lambertlab.us"]
    CF --> TS["Tailscale Mesh Ingress<br/>WireGuard Point-to-Point"]
    TS --> Traefik["Traefik Edge Ingress :443<br/>Wildcard TLSStore + Middlewares"]

    subgraph K3sCluster["K3s Hybrid Cluster (ArgoCD GitOps HA)"]
        Traefik --> Guac["Apache Guacamole Gateway"]
        Traefik --> KubeVirtMgr["KubeVirt Manager Web UI"]
        Traefik --> Catalog["30+ App Workloads"]

        subgraph L2VXLAN["Software-Defined L2 VXLAN Fabric (br-lab0: 10.10.0.0/24)"]
            Bridge["br-lab0 Virtual Bridge"]
            Guac -.->|"Multus net1: 10.10.0.50"| Bridge
            Bridge --> DC01["DC01 Active Directory<br/>Windows Server 2025: 10.10.0.10"]
            Bridge --> Win11["Win11 Admin Workstation<br/>Sysprep Specialized: 10.10.0.155"]
            Bridge <--> OPNsense["OPNsense Firewall VM<br/>Virtual Edge Routing"]
            CoreDNS["K3s CoreDNS<br/>health_check 5s max_fails 2"] -->|"Conditional Forward *.ad"| DC01
        end
    end

    subgraph StorageSAN["Centralized Storage (TerraMaster F4-425 Plus)"]
        DC01 -->|"Raw Block 80GB iSCSI LUN"| iSCSITarget[("SAN iSCSI Target LUNs")]
        Win11 -->|"Raw Block 64GB iSCSI LUN"| iSCSITarget
        Catalog -->|"Media & ISOs /Volume3/isos"| NFSPool[("14TB Bulk NFS Pool")]
    end

    subgraph Observability["Heavy Compute & Security Information and Event Management (SIEM)"]
        Logstash["Logstash Pipeline :12201 GELF"]
        Catalog -.->|"GELF over Tailscale"| Logstash
        Logstash --> ES[("Elasticsearch 8.x<br/>Local NVMe / ILM Single-Node")]
        ES --> Kibana["Kibana Dashboard :5601"]
        Traefik -->|"Ingress Route"| Kibana
    end

🚀 Core Architectural Highlights

  • Bare-Metal Virtualization: High-performance Windows Server 2025 and Windows 11 Enterprise LTSC virtual machines managed natively alongside containerized microservices via KubeVirt v1.9.0.
  • Hardware-Enforced Security: OVMF UEFI Secure Boot paired with persistent virtual TPM 2.0 (swtpm) state volumes backed by Longhorn.
  • Direct Line-Rate iSCSI Flashing: Template images are converted directly from compressed QCOW2 into raw iSCSI block LUNs via qemu-img convert (libiscsi), completely bypassing HTTP ingress proxy idle timeouts and Longhorn 64GB scratch volume overhead.
  • Automated Sysprep Specialization: Dynamic WIN-* machine naming, zero-touch OOBE regional bypass, and least-privilege domain join automation via declarative unattend.xml answer files.
  • Dual Directory Model: Seamless identity federation combining on-premises Active Directory (ad.lambertlab.us, NetBIOS LAMBERTLAB) with Microsoft Entra ID.
  • Entra Cloud Sync: Zero-touch Password Hash Synchronization (PHS) running under an Active Directory Group Managed Service Account (provAgentgMSA$).
  • Clientless Remote Gateway: Apache Guacamole provides browser-based HTML5 RDP/SSH access with Entra ID OIDC SSO, on-premises LDAPS authentication with custom Root CA JVM truststore injection, and native Multus L2 VXLAN pod networking.
  • NMState Multicast VXLAN (br-lab0): Multi-node Layer 2 broadcast overlay connecting physical nodes and virtual machines without requiring expensive 802.1Q managed switch VLAN trunking.
  • Direct Pod Bridging (Multus CNI): Apache Guacamole pod attaches a secondary interface (net1, 10.10.0.50/24) directly into br-lab0, communicating with domain controllers and desktops at wire speed with zero NAT overhead.
  • CoreDNS Conditional Health-Checking: In-cluster CoreDNS conditionally routes *.ad.lambertlab.us to DC01 (10.10.0.10) with automated health checks (health_check 5s, max_fails 2) preventing cluster-wide UDP timeout cascades.
  • Self-Healing Bridge Watchdog: Lightweight Kubernetes DaemonSet continuously detects and recovers un-enslaved VXLAN links caused by asynchronous USB NIC boot latency.
  • High-Availability GitOps: ArgoCD HA with Redis Sentinel, controller sharding, and 5 dedicated AppProjects reconciling 30+ services across deterministic synchronization waves (0 to 3).
  • Multi-Tiered Storage Matrix: Dedicated TerraMaster SAN iSCSI block LUNs for low-latency VM disks, distributed Longhorn block storage with synchronous cross-node replication for databases, 14TB high-capacity NFS pools for bulk media, and dedicated local NVMe storage for real-time SIEM indexing.
  • Centralized Log Ingestion (ELK Stack): Cross-node Docker container logs streamed directly to Logstash in the ELK Stack (Elasticsearch, Logstash, Kibana) via Docker GELF drivers over Tailscale on port 12201.
  • Strict Index Lifecycle Management (ILM) Retention Rule: Single-node Elasticsearch configurations enforce number_of_replicas: 0 across index templates to ensure green cluster health and prevent automated ILM rollover and pruning stalls.
  • Workstation Compute Isolation: Strict memory and CPU resource caps (guacd capped at 2 cores) preserve 30 Xeon threads and 20GB RTX compute for real-time Elasticsearch pipelines and machine learning workloads.

📖 Complete Documentation Index

🏛️ Architecture Deep-Dives

📦 Application Service Catalog

🛠️ Operational Runbooks