AI project
Aria: a sandboxed AI agent for every employee
One agent per person, sandboxed so nothing leaks between them.
Aria is how I configured OpenClaw for Genuka. Every employee has their own agent, and each agent is sandboxed so there is never a memory leak between people. I also set it up for Genuka's groups, where it helps manage the group, sends reminders and keeps track of the group's goals. It is not tied to a single model: I configured several LLMs, including GPT-4o and DeepSeek. I designed the architecture, deployed it, wrote the custom skills, and debugged the problems that only show up once real messages arrive.
This is a private deployment for Genuka, so there is no public link. Employee and client details are left out.
- Role
- Architecture, configuration, deployment and custom skills
- Stack
- OpenClawGPT-4oDeepSeekMulti-LLMWhatsAppDockerCoolifyNginx
Architecture
- One agent per employee, each sandboxed with its own memory and state, so there is never a memory leak from one person's agent to another's.
- Role-based permissions and admin oversight over what each agent can do.
- Individual personalities per agent, defined in plain files (SOUL.md, IDENTITY.md, USER.md) alongside a users.json registry and per-agent state files.
- Model-agnostic: several LLMs configured side by side (GPT-4o, DeepSeek and others), so the model can be chosen per need and cost.
- WhatsApp as the main channel.
Group agents
- Configured for Genuka's various groups to help manage each group.
- Sends reminders to the group.
- Keeps track of the group's goals.
Deployment
- Migrated from a local setup to Coolify on an Ubuntu VPS, running in Docker.
- Locked down who can talk to it with a pairing DM policy and an allow-list of numbers.
- Added a second WhatsApp account and a dedicated workspace to route messages from specific numbers to the right agent.
Skills I built
- Brain: a pre-call context manager that filters for relevance, compresses history and logs token usage, to cut token spend on every request.
- Coding assistant: a Planner, Developer and Reviewer pipeline. Work is routed to different models by task (a Sonnet model for bug fixes, Kimi K2 for features, Opus for review), it detects GitHub or GitLab automatically, and it has a hard rule against pushing to main or master. Agents exchange JSON contracts, the user gets explicit pause points, the auto-fix loop is capped, and the reviewer's severity ratings are calibrated.
- Humanize-writing: checks WhatsApp replies and outreach messages against Wikipedia's “Signs of AI writing” guide, so AI tells are caught before anything is sent.
Problems I solved in production
- WhatsApp settings reverting on restart, and 401 conflicts caused by stale session files.
- Session-takeover deadlocks on a local Windows instance, traced to multiple Node processes and a session-memory hook writing mid-process. Updating OpenClaw fixed it.
- Read-only sandbox volumes, fixed with a host-path bind mount.
- A main-agent deadlock triggered by replayed voice notes.
- Nginx WebSocket configuration that didn't survive container restarts.
- Agent workspace files being overwritten by accident across several agents.
Agents for real work
- An assistant for a business developer at Genuka that sharpens prospecting messages, tracks leads and keeps follow-ups on schedule.
- A lead-generation and WhatsApp outreach pipeline where Aria spawns a sub-agent to find and vet prospects.
What I'm exploring next
- Scheduled and event-triggered skills.
- Dynamic multi-agent spawning and shared scratchpads.
- Persistent memory with Mem0.
- Deeper tool integrations.