Rendered at 20:41:21 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
debazel 14 hours ago [-]
Why is it replacing true/false with T/F? true/false is already 1 token in all tokenizer I've seen. Even worse is replacing null with ∅. ∅ is a special unicode symbol that takes up 2 tokens compared to the 1 token for null...
cedws 13 hours ago [-]
Brand new GitHub account, brand new HN account. I stay far away from projects like this these days, they can easily be malicious. GitHub needs some kind of indicator for projects authored by tenured developers with a real identity.
plufz 11 hours ago [-]
And they need some kind of downvote. Projects needs to be able to lose a star.
fallingbananna 12 hours ago [-]
What's even weirder is that the substitution is explained in "How TOON works" section. Yet TOON never describes such behavior anywhere in its spec.
wehnsdaefflae 11 hours ago [-]
Maybe they want to be tokenizer agnostic? Then they would need to go by character count, right? Despite these inconsistencies, has anyone actually verified their promise? 350 vs. 10.000 tokens would still be very valuable even if they mess up some edge cases
wannabe44 11 hours ago [-]
Oh it is not returning full tool schema. You can't cut down that much by serialisation alone.
AmazingTurtle 14 hours ago [-]
↲ is also two tokens instead of a simple \n lmao
hnlmorg 12 hours ago [-]
How is an LF two tokens? Or were you referring to the Unicode symbol?
I took their example to mean an actual LF ASCII character but now Ive read your comment, maybe I was being too charitable?
I don’t think the Show Me section makes sense, the TOON variant clearly doesn’t have the same information. And the examples in the “How TOON works” section focuses on number of characters instead of tokens. I would think “null” is a single token anyway, why bother replacing it with an uncommon character?
sceptic123 12 hours ago [-]
Isn't there value to the verbose information too? Knowing what a tool does and what the inputs are increase the likelyhood of successful tool calls.
Loic 14 hours ago [-]
I spent more than one week, as a side project, to add an MCP server to my Cheméo website. Only 4 tools.
It took me way more time than expected, I was thinking: "Just wrap the REST API, 2h, done".
The MCP payload has nothing to do with the REST API one. Because you need to make it interpretable and context efficient even so it is structured data.
It was really interesting work and I suppose very little people are taking the time to rethink what is sent over the wire while creating a MCP server. If so, we would not have MCPs with the minimal payload being 500kB of JSON soup.
If you send my MCP through your "save token filter", I can guarantee you, that you will have trash down the line.
maxrev17 12 hours ago [-]
Yeah this is why a code execution sandbox so the ai can batch calls and select from the response format what it wants and limit the number of responses with instruction to be concise and preserve its context is a really cool thing to do.
spiderfarmer 12 hours ago [-]
This is where using a framework really shines. I used Laravel MCP which makes it trivial to add MCP tools to your CRUD.
ameshkov 14 hours ago [-]
I made an MCP proxy with a similar idea in the past: replace a ton of tools that consume tokens with just two (get_tool_schema, invoke_tool) - https://github.com/ameshkov/mcp-compress-router
One thing that I noticed is that it’s often better to return tool names with argument names, i.e. return “search_web(query)” instead of just “search_web” when listing tools. Otherwise models often tend to hallucinate argument names and an extra turn is required to correct the mistake.
One additional advantage that such tools provide is that when you use different coding agents you don’t have to set up all the MCP servers in every agent, you just set up one (or point the agent to the cli like in this project).
victor_edka 5 hours ago [-]
[dead]
moinism 13 hours ago [-]
How do unresearched, vibe-coded projects like this reach the front page?
eterm 12 hours ago [-]
I'm convinced that the majority of upvotes are based on reading a title rather than clicking through to an article.
People want a token efficient MCP CLI client. Whether this actually is one is less relevant.
wannabe44 11 hours ago [-]
I like to know the average age of accounts which upvoted this post.
maxrev17 12 hours ago [-]
Bots, bots everywhere
philipp-gayret 13 hours ago [-]
OP, I'm very interested in seeing an actual comparison ran through a common tokenizer of tool calls. I think you'll find different results than what you intended for this tool to be. You've mixed up tokens with characters on your screen.
stephantul 13 hours ago [-]
I think that some of these choices (as others have commented) show that the author has not investigated how tokenization works.
Tokenization is not some black box, you can run tokenizers and check them.
saretup 13 hours ago [-]
> zero information lost
You're just returning the name of the tool, the rest of the information (description/input schema) is definitely lost. Cut to the LLM making mistakes in calling the tool with incorrect schema or calling the wrong tools altogether, recovering, wasting tokens and cycles.
anshumankmr 9 hours ago [-]
I like the idea, but this seems a little too aggressive, JSON (287 tokens) — what every other MCP client returns:
~~~
[
{"name": "search_web", "description": "Search the web for information",
"inputSchema": {"type": "object", "properties": {"query": {"type": "string", "description": "Search query"}, "num_results": {"type": "number", "default": 5}}, "required": ["query"]}},
{"name": "fetch_url", "description": "Fetch content from a URL",
"inputSchema": {"type": "object", "properties": {"url": {"type": "string"}}, "required": ["url"]}}
]
TOON (5 tokens) — what mcptoon returns:
search_web fetch_url
~~~
4 hours ago [-]
wannabe44 14 hours ago [-]
I am not going to trust a single number thrown by these AI hustlers written in that salesman voice.
Leave alone 97%.
> Your agent calls 20 tools. Each returns 500-3,000 tokens wrapped in {"content":[{"type":"text","text":"..."}]}.
This is a problem with your tool design. Most MCPs are fully vibe coded without any thought about tool selection.
> On a 128K context window, that's 30-55% gone. Not on work. On syntax.
Tool output is not "syntax" you donkey clanker.
Again, use the code approach, let the LLM filter out the JSON using tools. This TOON thing is just vibes. Most of the time your tool output should not even be JSON. It should be well formatted markdown. In cases where it's large structured data, your LLM should have tools (code / jq) to dissect it. So TOON is pointless.
liminal-dev 10 hours ago [-]
I’m going to start using “donkey clanker”.
alxhslm 14 hours ago [-]
Don’t quite see the point of this. It is well known that MCP is a bit bloated for coding agents at least.
But, why not just use CLIs for each tool? That seems to be where things are going anyway
And using MCP as an internal communication method seems odd when you could use the APIs directly
codingjoe 8 hours ago [-]
Q: aren't models trained to message templates using JSON for tool calls. Would a model inherently struggle with a different format?
Q: is there a measurable difference compared to harnesses with tool search?
dthedavid 15 hours ago [-]
How does it work? Im building a video editor and right now it has access to nearly 100 tools. Would be good to learn the techniques you used to make tool discovery more efficient.
arjie 15 hours ago [-]
The readme has some examples for what it does. It doesn’t list the entire schema (noisy). Instead it uses shorthand. Perhaps a sufficiently smart agent can do this.
kk3838368397373 14 hours ago [-]
sorry, is Headroom still a thing? What happened to it? Is anyone still using it? so many things , which one is actually working :/ idk this ai world
Are they though? Will they always be? Is it in their interest to be efficient?
ekisu 9 hours ago [-]
Somewhat related to this project, I'm surprised that not all harnesses are using something like CodeMode for MCPs.
Been experimenting with it in the OpenCode V2 beta and it's pretty great. The combination of tool search, call chaining and field projections feels just right and saves a lot of context. LLMs are good at writing code, who would have thought that?
bythreads 15 hours ago [-]
Sorry, isnt this just compression? Lookups burn tokens just on the other end?
hnlmorg 14 hours ago [-]
I really think we’ve missed a trick using JSON instead of SExpressions as the default marshaller for AI tool use.
Avery29 11 hours ago [-]
Making MCP context cost visible before the agent sees it feels like a useful debugging tool, not just an optimization.
vasco 15 hours ago [-]
I really doubt that null and \n make any sense to replace with non ascii symbols. They are both most likely already a token only and for other purposes at least \n becomes larger as a symbol.
I took their example to mean an actual LF ASCII character but now Ive read your comment, maybe I was being too charitable?
https://github.com/activeing123/mcptoon/blob/main/src/mcptoo...
It took me way more time than expected, I was thinking: "Just wrap the REST API, 2h, done".
The MCP payload has nothing to do with the REST API one. Because you need to make it interpretable and context efficient even so it is structured data.
It was really interesting work and I suppose very little people are taking the time to rethink what is sent over the wire while creating a MCP server. If so, we would not have MCPs with the minimal payload being 500kB of JSON soup.
If you send my MCP through your "save token filter", I can guarantee you, that you will have trash down the line.
One thing that I noticed is that it’s often better to return tool names with argument names, i.e. return “search_web(query)” instead of just “search_web” when listing tools. Otherwise models often tend to hallucinate argument names and an extra turn is required to correct the mistake.
One additional advantage that such tools provide is that when you use different coding agents you don’t have to set up all the MCP servers in every agent, you just set up one (or point the agent to the cli like in this project).
People want a token efficient MCP CLI client. Whether this actually is one is less relevant.
Tokenization is not some black box, you can run tokenizers and check them.
You're just returning the name of the tool, the rest of the information (description/input schema) is definitely lost. Cut to the LLM making mistakes in calling the tool with incorrect schema or calling the wrong tools altogether, recovering, wasting tokens and cycles.
search_web fetch_url
~~~
Leave alone 97%.
> Your agent calls 20 tools. Each returns 500-3,000 tokens wrapped in {"content":[{"type":"text","text":"..."}]}.
This is a problem with your tool design. Most MCPs are fully vibe coded without any thought about tool selection.
> On a 128K context window, that's 30-55% gone. Not on work. On syntax.
Tool output is not "syntax" you donkey clanker.
Again, use the code approach, let the LLM filter out the JSON using tools. This TOON thing is just vibes. Most of the time your tool output should not even be JSON. It should be well formatted markdown. In cases where it's large structured data, your LLM should have tools (code / jq) to dissect it. So TOON is pointless.
But, why not just use CLIs for each tool? That seems to be where things are going anyway
And using MCP as an internal communication method seems odd when you could use the APIs directly
Q: is there a measurable difference compared to harnesses with tool search?
Been experimenting with it in the OpenCode V2 beta and it's pretty great. The combination of tool search, call chaining and field projections feels just right and saves a lot of context. LLMs are good at writing code, who would have thought that?