Local + Cloud

Two local servers, connected through one governed AI fabric.

InnerChispa is built to reduce dependency on external credits without becoming isolated from the best models. Local AMD and Intel nodes, RalphiIA MCP, Mongo coordination, Open WebUI, Cloudflare, Google Cloud, Alibaba Cloud and external model APIs each have a clear role.

Two local servers, connected through one governed AI fabric.
Generated and reused visual assets are composed as editorial website imagery; diagrams remain readable HTML.

Server Fabric

Two local nodes, one coordination layer.

The public site describes roles without exposing private IPs or credentials.

01

Primary Intel Node

Coordination services, browser/build/test work, fallback execution, MCP access, Open WebUI integration and daily operational services.

02

AMD Local AI Node

Local-first AI capacity for coding, heavy reasoning and private model experiments, with vLLM/ROCm direction and rollback-aware model registry.

03

RalphiIA MCP

Shared bridge for ChatGPT, Codex, Cursor, Open WebUI and agents. It exposes tools, state, runbooks, documents and operational actions under governance.

04

Cloud + Edge

Cloud Run, Cloudflare, Google Cloud, Alibaba Cloud and external model providers extend reach when public access, scale or advanced models are needed.

Primary local server

The primary local environment carries coordination services, browser/build/test work, operational APIs, MCP access, Open WebUI integration and day-to-day services.

AMD local AI server

The AMD node is the strategic local AI brain for heavier coding and reasoning workloads, local model experiments, private inference and future vLLM/ROCm routing.

MCP interconnection

RalphiIA MCP is the bridge: it lets ChatGPT, Codex, Cursor, Open WebUI and agents share tools, state, runbooks and operational actions.

Credit savings

Local-first execution avoids spending cloud credits on every coding, reasoning, summarization or automation task. The system should know when local is enough.

Cloud escalation

When a task needs stronger models, public hosting, serverless scale, crawler data, or provider-specific tools, the system can escalate to OpenAI, Gemini, Claude, Qwen providers, Google Cloud, Cloudflare, Alibaba Cloud, Bright Data or other services.

Evidence and fallback

Model routing should expose selected node, selected model, backend, reason, health and fallback status so operators can see where work actually ran.

Architecture Diagram

From conversation to governed action.

01Human / Voice / UI

Operators ask, approve and supervise.

02RalphiIA

Persistent operational intelligence.

03InnerOS

Memory, agents, models, tools, tasks, guards and devices.

04Local + Cloud Compute

Local-first when control matters; cloud when capability matters.

05Business Systems

Quotes, workforce, field service, security, documents and evidence.

AMDIntelOpenAIGoogle CloudGeminiQwenClaudeDockerMongoDBGitHubCloudflareSupabaseReactTypeScriptPythonFastAPIMCPWhatsAppNotionZKTecoHikvisionUiPathDevpostlablab.ai