Rendered at 10:53:43 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
shostack 20 hours ago [-]
Love the post and have a similar setup. But I have questions about how you handle sensitive personal info with Codex and Claude with this setup given Hermes has no boundaries on its own with its integrations with Codex for example. It can leak memory and tool calls and context through.
At home I have Hermes on a VPS with matrix and mnemosyne and it is largely for personal use. While I was excited to try out the seamless Codex wrapper, this leakage was unacceptable compared to a ZDR OpenRouter provider like fireworks.
Where I struggle is now in architecting a clean way to have Hermes separate out personal stuff from coding stuff, and then orchestrate coding through Codex as that seems to be the only way to use it headlessly on my phone since they have no way of doing that in the ChatGPT app right now on Linux.
Does Buzz let you neatly silo that?
I'm struggling between:
"I want to use Hermes with frontier models with a seamless wrapper on my $20 ChatGPT pro subscription so I can go through my Hermes setup for everything including coding"
And
"OpenAI and Anthropic are not ZDR and I am not comfortable sending sensitive personal and family data to them. But I at least want a mobile-friendly remote coding experience with them on my phone and VPS"
carimura 11 hours ago [-]
It depends on what you mean by sensitive personal data. I actually don't give any agents access to my main email boxes instead they access a business google workspace that is spinning up.
I do separate the agents duties somewhat to try and give them the least privileges they need to do a job, and talk to all of them through buzz, and they talk to each other (occasionally) when they need something they don't have access to.
As mentioned in the post I still do a majority of my coding at my local terminal with just bits and pieces sent to my "dev-agent" through Buzz. This is a next experiment for me - (that's a hyphen!) to offload more coding to agents.
I already have that effectively via pinning Fireworks on Openrouter. I am unclear how this would solve the architecture issue I raised though.
whazor 2 days ago [-]
I am now on the MCP route, where I have lots of personally hosted MCPs. Including managing calender/e-mail.
It does not cost too much effort to maintain MCP servers. No port forwarding or VPNs required thanks to OpenAI tunnels.
And security wise its quite nice, since you have to activate MCP or give permission sometimes. So each chat is kind of isolated from each-other.
carimura 2 days ago [-]
interesting.. does that bloat the context window? I imagine Hermes already struggles with context window size.. been meaning to look into that.
trollbridge 2 days ago [-]
Couple of notes on this
First, it matters how well the MCP is written; a well-written one won't be so massive
Secondly, different sessions should load different MCPs. Load only what is needed
Thirdly, if you're finding you need a dozen MCPs, you need to consolidate that into a single service that does the work and then presents a single, unified API (via MCP) to your agent
Finally, start using models with 1M context window limits. I really don't know how people can stand being stuck at a limit like 272K.
carimura 2 days ago [-]
ya all makes sense. hoping the harness can manage some of the MCP noise. I'm using 5.6-Sol so context window is 1M.
cyanydeez 2 days ago [-]
Do you not use any context pruning?
Atotalnoob 2 days ago [-]
Skills and MCPs CAN bloat the context window. Some harnesses do more progressive loading of skills/mcp depending on how many you have. Check your harnesses docs.
Personally, I feel skills+CLI is better, since only the description of the skill enters the context window until it’s required and CLIs should support help flags which will allow you do have progressive discovery.
And a CLI is also easy to use yourself.
whazor 2 days ago [-]
I don't have problems on ChatGPT with the amount of MCPs (five) I have. But each MCP can do quite a lot.
orangebread 2 days ago [-]
Very cool setup. I went down a similar path and while it is definitely cool to see the extent of how agentic AI can be, I agree the ROI isn't _quite_ there yet.
I don't live in a world where I need to be constantly reading and replying to emails, so 95% of my inbox is just subscription spam. With development I need to be at the helm to design the planning requirements and actively make decisions before the agent goes off and executes the plan. But I don't need a personal agent for that, I work directly out of codex/claude-code.
carimura 2 days ago [-]
same. I'd like to get to a place where I can have long-running tasks and goals for more over-night working, but it's not there yet.
lw18511811620 2 days ago [-]
I've landed on a similar conclusion building a much narrower tool: for messaging specifically, I don't think people want an agent that reads and sends on their behalf — they want the drafting to be fast, not the sending to be automatic. The moment you take the human out of the send button, the failure modes (wrong tone, made-up facts) become a lot more costly. Scoping down to "give me two good drafts, I'll pick and edit" turned out to be more useful day-to-day than anything closer to full autonomy.
kolli_kumar 16 hours ago [-]
This is awesome! I have a similar setup, with multiple Nanobot Docker instances across three different VMs, along with Matrix and Gitea. I also have a few connected with Obsidian.
I’d love to understand more about your Obsidian setup and Buzz.
The Matrix app on iOS isn’t as good as Telegram.
I have 16 Nanobot instances—4 for my parents and 12 for me and my spouse.
With OpenWebUI and its Knowledge feature, I was thinking of moving some of them there.
I have only one small VPS for ntfy and Matrix. Everything else is self-hosted.
A few years ago, I bought way too much RAM for my Dell server and added more storage just to experiment. Now I’m sitting on a gold mine, with prices having nearly tripled!
I also have a Synology for VM snapshots and backups to Google Drive and B2.
carimura 11 hours ago [-]
Obsidian is like the shared knowledge base that everyone has access to and that I can easily read/review (although I try not to too often). It has architecture stuff, go-to-market section, howtos (setting up new agents, dns backups, etc.), info on my main projects, writing, etc. It gets synced between my machine and the VPS. I also store in github as a backup occasionally (thx for reminder!).
I use the actual Buzz ios app which I was building on my own but at first glance it appears is now in the iOS store.
richardkam0511 2 days ago [-]
I actually went down a different path and set it up so my friend and I can use the same agent with approval-gated turns. He's more technical than I am, but I believe I'm better at marketing. We can see both of our prompts and be on the same page. There is also a mode where we can branch off with our own agents.
Watching his prompts made me better at producing my own and using the agent effectively.
carimura 2 days ago [-]
That's neat. You could do this with Buzz by just having open room conversations with agents.
richardkam0511 2 days ago [-]
I checked it out, looks pretty cool. The tool I built is pretty similar, usepoly.co if you want to take a peek!
moribvndvs 2 days ago [-]
Not pointed at the author, but at the current state of affairs: this is fucking exhausting. We went and made a trillions-dollar market out of the bikeshedding maximization machine.
morkalork 2 days ago [-]
With Claude not only can you shave more yaks, you can build and run a yak farm and specialized yak-trimmer factory!
carimura 2 days ago [-]
this author agrees. hence my attempt at steering part of that trillion dollars for my non-profit which maybe someday might do some good.
Manfrednotfunny 2 days ago [-]
No we did not.
LLM was invented and its just clear that this is something someone needs to build.
Why?
Because it makes just sense. You don't want an agent running on a laptop you close. You want to keep context small, you want to split up work / parallize it etc.
I'm now waiting for a while until the open source agent platform emerges and im borderline motivated to build something but i'm not doing it. He did, which is not a crime.
moss_dog 15 hours ago [-]
Thanks for the write-up! I'm interested in doing something similar. How are you sandboxing the agents?
carimura 11 hours ago [-]
I use the term sandbox loosely. They're on a VPS running as user hermes but they share everything with each other in their own user space. I could/should probably improve this. Maybe make them their own unix user. They are separated from the root-access agent.
lukasco 2 days ago [-]
I've been building a triage agent for my inbox and whatsapp (it's product shaped), which has ironically left me not building one of these. So even while productizing, I'm getting fomo on the full monty.
I've also been building a harness that maintains my apps which I'm hoping to open source.
Hard agree that these things don't have personal ROI, and are actually quite hard to build reliably.
But it's really fun! And having a bot fix a live error is pretty exciting.
owencmcgrath 2 days ago [-]
The most interesting part of this to me is having an agent that reads logs and spins up fixes in real time.
Great write up!
jessepcc 1 days ago [-]
I put my setup in a local machine for easier use of browser.
How agents know other's capabilities? Shout out and call for help? Is it handled by Buzz?
carimura 1 days ago [-]
buzz is how they communicate. they know about each other through a couple of ways: their own stored memory, a Mnemosyne bank, or a shared obsidian wiki. Or maybe they joined a linkedin for bots that I don't know about.
dizhn 23 hours ago [-]
Paseo agents can use a browser on the desktop app. You can be connected to local or remote agents. It doesn't matter.
avereveard 2 days ago [-]
tried similar setups in the past, but messages is just not my jam, so built a html wrapper on top of claude code and open code that runs on a small-ish instance. looks like this https://i.imgur.com/lj9Fgco.png (screenshot anonymized with chatgpt) and allows multiple conversation per project, and has a few conveniences like scheduled tasks. one of the project in the list is the project itself, so I can add features whenever.
carimura 2 days ago [-]
Ya I'm starting to see that agents are just "input + llm + some memory + tools + lots of cron jobs".
tln 2 days ago [-]
That looks really useful
You haven't put this on github have you?
avereveard 2 days ago [-]
No it is a vibecoded mess. Bet if you show the screen to a agent of your choice you will have yours in no time. The benefit is that if you give it self installing scripts you can just ask itself to add features
Cgroup and nice it so you never lose control of the machine tho
aliasxneo 2 days ago [-]
Have you thought about browser/machine use at all? I like the idea of Grok Bot but would rather host/maintain it myself with my own agents.
carimura 2 days ago [-]
Not yet. The agents do have access to the web but not really a proper browser. Browserbase or equivalent is next on my list.
giwook 1 days ago [-]
Out of curiosity why Buzz and not Discord if you're looking for a Slack alternative?
carimura 1 days ago [-]
two reasons that I am moderately confident in: 1) I really want something open source, but most importantly, 2) I want something easy.
My experience trying to get Discord and/or Slack to work was painful for 1 agent, let alone the potential for hundreds.
Here's the perfect example of the power of Buzz: there was some hype today about Grok Bot so i wanted to spin it up. I did, and within like 3 minutes, I had a Grok Bot talking to my other agents in Buzz, and Buzz is definitely not a supported plugin.
slowhogs 1 days ago [-]
[dead]
noashavit 19 hours ago [-]
This is a great post, thanks for sharing!
JSR_FDED 2 days ago [-]
Appreciate the honesty and lack of hype.
tosh 2 days ago [-]
ty for the writeup!
being able to talk to each of the agents via dm (but also in group chats) sounds interesting
does that mean that you have 1 chat per domain specific agent? can you also start multiple sessions/threads or is that not part of the way you interact with them currently?
carimura 2 days ago [-]
it's just like Slack really. I have a DM open for each agent where I can talk direct. But they are also part of individual rooms also where I just need to @ them and they open a thread. They can talk to each other also by @'ing each other (in public rooms or ones they both belong to)
It's literally just like humans.. except. not.
tosh 2 days ago [-]
makes sense, ty
4lx87 2 days ago [-]
Have you experimented with using a single executive agent (instead of multiple agents)?
carimura 2 days ago [-]
that was one of my FAQs at the bottom. I want separation of duties and least-privilege so agents can only access what they need for their particular duties.
That said, my vision is to eventually build manager agents to manage the minion agents. Like real people. I have no idea if "people" is the right analogy for all this work but my brain can't really wrap around a different analogy yet.
gchp 2 days ago [-]
Cool! How much does this setup cost you, roughly?
carimura 2 days ago [-]
$48/month for the droplet (could probably be cheaper on Hertzner) and $100/month for OpenAI plan. so ~$150/month but again this could be cheaper with a different VPS and using Terra/Luna.
I also have a claude max plan for coding but like I mentioned in the article coding is still separate from the agents.
messh 2 days ago [-]
A box with 2vcpu, 4gb ram and 50gb hdd on https://shellbox.dev is $0.02/hr and you pay per minute only for what you use. Even using it 24/7 is like less than $15
carimura 2 days ago [-]
ya i'm sure there are lower cost options. My droplet is 8 gigs mem which I found necessary for 6 agents.
messh 20 hours ago [-]
The x2 instance has 8gb memory. They go up to x8. And if you don't use it say during a weekend, then you stop it and resume on Monday, not paying for that period
Cameri 1 days ago [-]
Nice to see Nostr show up randomly!
carimura 1 days ago [-]
that's exactly what I thought too. When I first saw Buzz I was like "oh, finally a reason to look at Nostr."
geooff_ 2 days ago [-]
It's always fun seeing how other people are using these tools.
> Has it been worth it? For the journey, yes, for the ROI, nope.
It's also nice seeing someone experiment without succumbing to AI psychosis.
carimura 2 days ago [-]
i'd like to go into the psychosis but i'm not there yet.
davidw 2 days ago [-]
What is the total monthly spend, approximately, on all this; if you don't mind saying, of course?
carimura 2 days ago [-]
(copied from comment above)
$48/month for the droplet (could probably be cheaper on Hertzner) and $100/month for OpenAI plan. so ~$150/month but again this could be cheaper with a different VPS and using Terra/Luna.
I also have a claude max plan for coding but like I mentioned in the article coding is still separate from the agents.
world2vec 2 days ago [-]
This is an ad
carimura 2 days ago [-]
that's news to me. for what?
johntash 1 days ago [-]
To subscribe to your rss feed, of course.
\s but it made me subscribe, it's hard finding ai-related content that isn't llm-generated and also isn't crazy people.
carimura 1 days ago [-]
ah forgot about the RSS feed. thanks for the complement!
phplovesong 2 days ago [-]
I would never have the power to have AI just do everything. I like to create, not have some slop machine do whatever.
sublinear 2 days ago [-]
I have gone back and forth on a similar sentiment a few times. I settled on this: the only responsible use of LLMs is to fill in gaps.
It's fine as a search tool or autocomplete. It can be okay to generate code if you didn't know how else to get started, or you've already limited the damage it would do by your own design.
People who overuse LLMs are usually trying to compensate for their lack of experience, structure in their work, or dysfunctional teams. Anyone trying to get hired should recognize it as a new red flag attempting to cover up the old red flags.
carimura 2 days ago [-]
AI is far from doing everything in my setup. just looking for incremental wins to see if my slop machine can be useful.
phplovesong 5 hours ago [-]
These days it seems that just having AI dom SOMETHING is the goal, no one cares about what it does, or what kind of quality it runs.
Its basically just a weird limbo thing that runs and you dont care how bad it is, not even considering what bugs it has. Outout is usually only glanced over.
At home I have Hermes on a VPS with matrix and mnemosyne and it is largely for personal use. While I was excited to try out the seamless Codex wrapper, this leakage was unacceptable compared to a ZDR OpenRouter provider like fireworks.
Where I struggle is now in architecting a clean way to have Hermes separate out personal stuff from coding stuff, and then orchestrate coding through Codex as that seems to be the only way to use it headlessly on my phone since they have no way of doing that in the ChatGPT app right now on Linux.
Does Buzz let you neatly silo that?
I'm struggling between: "I want to use Hermes with frontier models with a seamless wrapper on my $20 ChatGPT pro subscription so I can go through my Hermes setup for everything including coding"
And
"OpenAI and Anthropic are not ZDR and I am not comfortable sending sensitive personal and family data to them. But I at least want a mobile-friendly remote coding experience with them on my phone and VPS"
I do separate the agents duties somewhat to try and give them the least privileges they need to do a job, and talk to all of them through buzz, and they talk to each other (occasionally) when they need something they don't have access to.
As mentioned in the post I still do a majority of my coding at my local terminal with just bits and pieces sent to my "dev-agent" through Buzz. This is a next experiment for me - (that's a hyphen!) to offload more coding to agents.
It does not cost too much effort to maintain MCP servers. No port forwarding or VPNs required thanks to OpenAI tunnels.
And security wise its quite nice, since you have to activate MCP or give permission sometimes. So each chat is kind of isolated from each-other.
First, it matters how well the MCP is written; a well-written one won't be so massive
Secondly, different sessions should load different MCPs. Load only what is needed
Thirdly, if you're finding you need a dozen MCPs, you need to consolidate that into a single service that does the work and then presents a single, unified API (via MCP) to your agent
Finally, start using models with 1M context window limits. I really don't know how people can stand being stuck at a limit like 272K.
Personally, I feel skills+CLI is better, since only the description of the skill enters the context window until it’s required and CLIs should support help flags which will allow you do have progressive discovery.
And a CLI is also easy to use yourself.
I don't live in a world where I need to be constantly reading and replying to emails, so 95% of my inbox is just subscription spam. With development I need to be at the helm to design the planning requirements and actively make decisions before the agent goes off and executes the plan. But I don't need a personal agent for that, I work directly out of codex/claude-code.
I’d love to understand more about your Obsidian setup and Buzz.
The Matrix app on iOS isn’t as good as Telegram.
I have 16 Nanobot instances—4 for my parents and 12 for me and my spouse.
With OpenWebUI and its Knowledge feature, I was thinking of moving some of them there.
I have only one small VPS for ntfy and Matrix. Everything else is self-hosted.
A few years ago, I bought way too much RAM for my Dell server and added more storage just to experiment. Now I’m sitting on a gold mine, with prices having nearly tripled!
I also have a Synology for VM snapshots and backups to Google Drive and B2.
I use the actual Buzz ios app which I was building on my own but at first glance it appears is now in the iOS store.
LLM was invented and its just clear that this is something someone needs to build.
Why?
Because it makes just sense. You don't want an agent running on a laptop you close. You want to keep context small, you want to split up work / parallize it etc.
I'm now waiting for a while until the open source agent platform emerges and im borderline motivated to build something but i'm not doing it. He did, which is not a crime.
I've also been building a harness that maintains my apps which I'm hoping to open source.
Hard agree that these things don't have personal ROI, and are actually quite hard to build reliably.
But it's really fun! And having a bot fix a live error is pretty exciting.
Great write up!
How agents know other's capabilities? Shout out and call for help? Is it handled by Buzz?
You haven't put this on github have you?
Cgroup and nice it so you never lose control of the machine tho
My experience trying to get Discord and/or Slack to work was painful for 1 agent, let alone the potential for hundreds.
Here's the perfect example of the power of Buzz: there was some hype today about Grok Bot so i wanted to spin it up. I did, and within like 3 minutes, I had a Grok Bot talking to my other agents in Buzz, and Buzz is definitely not a supported plugin.
being able to talk to each of the agents via dm (but also in group chats) sounds interesting
does that mean that you have 1 chat per domain specific agent? can you also start multiple sessions/threads or is that not part of the way you interact with them currently?
It's literally just like humans.. except. not.
That said, my vision is to eventually build manager agents to manage the minion agents. Like real people. I have no idea if "people" is the right analogy for all this work but my brain can't really wrap around a different analogy yet.
I also have a claude max plan for coding but like I mentioned in the article coding is still separate from the agents.
> Has it been worth it? For the journey, yes, for the ROI, nope.
It's also nice seeing someone experiment without succumbing to AI psychosis.
$48/month for the droplet (could probably be cheaper on Hertzner) and $100/month for OpenAI plan. so ~$150/month but again this could be cheaper with a different VPS and using Terra/Luna.
I also have a claude max plan for coding but like I mentioned in the article coding is still separate from the agents.
\s but it made me subscribe, it's hard finding ai-related content that isn't llm-generated and also isn't crazy people.
It's fine as a search tool or autocomplete. It can be okay to generate code if you didn't know how else to get started, or you've already limited the damage it would do by your own design.
People who overuse LLMs are usually trying to compensate for their lack of experience, structure in their work, or dysfunctional teams. Anyone trying to get hired should recognize it as a new red flag attempting to cover up the old red flags.
Its basically just a weird limbo thing that runs and you dont care how bad it is, not even considering what bugs it has. Outout is usually only glanced over.