Stop Wasting LLM Tokens! I Built a Rust CLI to Prune JS/TS Codebases by 80% ๐ฆ๐
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ THE โINFINITE CONTEXTโ TRAP โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค โ 1. Attention Degradation โ Lost-in-the-Middle: critical interfaces get โ โ โ buried under repetitive DOM noise and loops. โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค โ 2. KV-Cache Prefill Lag โ Time-to-First-Token (TTFT) scales with promptโ โ โ size; 150k+ raw tokens stall your agent. โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค โ 3. The โTailwind Taxโ โ Paying frontier API rates to ingest 80-char โ โ โ strings like โflex items-center justify-โฆโ โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค โ 4. Rate-Limit Throttling โ Bloated prompts quickly exhaust TPM (Tokens โ โ โ Per Minute) quotas in CI/CD pipelines. โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Have you ever dumped an entire React or Next.js repository into Claude 3.5 Sonnet, GPT-4o, or a local Ollama model to ask: โHow does authentication state flow through my UI, and what endpoints handle it?โ If you inspect the prompt you sent, over 70% of the tokens are dead weight: Hundreds of lines of static Tailwind CSS utility classes (className=โflex flex-col items-center justify-between p-8 bg-white dark:bg-zinc-950 rounded-2xl shadow-xlโฆโ). Imperative loops, array mappers, and string formatting logic that have nothing to do with application architecture. Missing birdโs-eye views: no clear backend API route table, no dependency graph, and no clean component hierarchy. While building an agentic Chrome extension powered by local LLMs, my context window collapsed: 209,757 tokens per scan. Responses took forever, local inference crawled, and the model routinely hallucinated core functions because key architectural interfaces were buried under syntactic noise. I built urai-ecma: a multi-threaded CLI tool written in Rust that uses SWC (Speedy Web Compiler) to parse JavaScript and TypeScript into Abstract Syntax Trees (AST). Instead of blindly concatenating files together like a text scraper, it acts as a semantic compiler for prompt engineeringโcompressing that same 209k token codebase down to 36k tokens (an 82.7% reduction) in milliseconds. Here is how it works, how it is architected under the hood, real benchmarks, and the engineering trade-offs you should know before using it. ๐๏ธ The Philosophy: What is โUraiโ? In classical Tamil literary heritage, monumental masterworks like the Thirukkuแนaแธท (เฎคเฎฟเฎฐเฏเฎเฏเฎเฏเฎฑเฎณเฏ) and Tolkฤppiyam (เฎคเฏเฎฒเฏเฎเฎพเฎชเฏเฎชเฎฟเฎฏเฎฎเฏ) contained dense, multi-layered philosophical thought. To make these works practical without destroying their architectural depth, classical scholars practiced เฎเฎฐเฏ เฎเฎดเฏเฎคเฏเฎคเฎฒเฏ (Urai Ezhuthudhal). Master commentators (Uraiyฤsiriyars) like Parimelazhagar and Ilampuranar did not just copy or mechanically summarize texts. They performed structural distillation: Isolating the core semantic axioms of each stanza. Stripping linguistic ornamentation that obscured meaning. Exposing grammar, intent, and relationships for reasoned debate. Modern enterprise JavaScript and TypeScript codebases are the epic literatures of software engineering. When asking an LLM to reason about your code, it doesnโt need raw syntactic exhaustionโit needs the structural anatomy, API contracts, state flows, and component signatures. urai-ecma acts as a modern Uraiyฤsiriyar for your codebase. โก Synthesis vs. Blind Concatenation Tools like repomix, gitingest, and code2prompt are file dumpers. They walk your directory, wrap raw text in XML/Markdown fences, and pass every single line of styling directly into your modelโs context. urai-ecma is an AST-aware compiler engine. Rather than treating code as raw strings, it parses your source into concrete syntax trees using ByteDance/Vercelโs swc_ecma engine and applies deterministic, semantic transformations: โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ URAI COMPILER PIPELINE โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Enterprise Monorepo (.ts, .tsx, .js, .mjs, .json) โ โผ [ignore::WalkBuilder (Rust)] Honor .gitignore, prune node_modules & dist โ โผ [Rayon Parallel Work-Stealing] Multi-threaded AST parsing across all CPU cores โ โโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโ โผ โผ [swc_ecma_parser] [swc_ecma_parser] Worker Thread A Worker Thread B โ โ โโโบ [RouteVisitor] โโโบ [RouteVisitor] โ Next.js/Express/NestJS โ Next.js/Express/NestJS โ โ โโโบ [ReactComponentAnalyzer] โโโบ [ReactComponentAnalyzer] โ Props, State, Hooks, JSX โ Props, State, Hooks, JSX โ โ โโโบ [ReactJsxPruner] โโโบ [ReactJsxPruner] โ Tailwind static class strip โ Tailwind static class strip โ โ โโโบ [FunctionSummarizerVisitor] โโโบ [FunctionSummarizerVisitor] Preserve structural stubs Preserve structural stubs โ โผ [Foyer Hybrid Cache (Disk + RAM)] Sha512_256 + Zstd compression โ โผ [swc_ecma_codegen + Tiktoken Engine] Emits high-density Markdown prompt + BPE o200k report ๐ฏ The Three Pillars of AST Token Optimization 1. AST Structural Stubbing (is_structural_stub_stmt) Traditional minification forces a bad compromise: either include full function bodies (wasting thousands of tokens on loops and math) or strip functions down to empty signatures (which deletes hooks, event listeners, and JSX layouts). urai-ecma solves this through Structural Stubbing. It inspects AST statements and retains only nodes critical to architectural comprehension: // Only statements defining component anatomy are preserved: fn is_structural_stub_stmt(stmt: &Stmt) -> bool { match stmt { Stmt::Decl(Decl::Fn()) => true, // Nested helper declarations Stmt::Decl(Decl::Var(var_decl)) => var_decl.decls.iter().any(|decl| { if let Some(init) = &decl.init { matches!(**init, Expr::Arrow() | Expr::Fn()) } else { false } }), Stmt::Expr(expr_stmt) => { if let Expr::Call(call_expr) = &*expr_stmt.expr && let Callee::Expr(callee_expr) = &call_expr.callee && let Expr::Ident(ident) = &**callee_expr { let name = ident.sym.as_ref(); // Preserves React Hooks, lifecycle timers, and global listeners: return name.starts_with(โuseโ) || name == โsetTimeoutโ || name == โsetIntervalโ || name.contains(โaddEventListenerโ) || name.contains(โrequestIdleCallbackโ); } false } Stmt::Return(ret_stmt) => { // Preserves JSX layout hierarchies: if let Some(arg) = &ret_stmt.arg { matches!( &**arg, Expr::JSXElement() | Expr::JSXFragment() | Expr::Paren() ) } else { false } } _ => false, // Computational loops, arithmetic, & validations are pruned } } Hooks stay visible: useEffect(() => { โฆ }, [dep]) remains intact, signaling side-effects to the LLM. JSX hierarchies stay visible: The complete return layout tree remains legible. Imperative logic is replaced: Loops, arithmetic, and validation are replaced with a single synthetic docstring comment. 2. Dynamic-Aware Tailwind Pruning Modern utility CSS accounts for massive token bloat. urai-ecma provides 4 modes (remove, remove_aggr, summarize, preserve): Preserves Dynamic Expressions: If you use dynamic styling like className={clsx(โbtnโ, isActive && โbtn-activeโ)} or ternary conditions, they are kept 100% intact. Strips or Summarizes Static Strings: Long static class lists exceeding the threshold (default: 96 chars) are either cleanly stripped or sent to a local Ollama model to produce natural-language style summaries (e.g., /* UI: Frosted glass card with dark mode /). 3. JSDoc-First Resolution with Local Ollama Fallback Summarizing every single function with an LLM is slow. urai-ecma uses a two-tier resolution strategy: JSDoc First: It checks whether human-written JSDoc annotations (@description, @param, @return) already exist. It even includes a proximity-scan fallback (within a 300-byte span) to associate detached comments. This takes 0 milliseconds of LLM compute! Local Ollama Fallback: If a function exceeds your line threshold (default: 5 lines) and has no JSDoc, it queries a local Ollama model (gemma4, llama3.2). No code ever leaves your machine. Foyer Hybrid Caching: Results are stored in a dual-tier cache powered by Rustโs foyer crate (64MB direct RAM buffer + 128MB Zstd-compressed disk storage with Sha512_256 keys). ๐ Real-World Before & After Look at what happens to a bloated React component when passed through urai-ecma: โ The Bloated Source Code (Before) const ErrorUI = ({ headerDescTxt = โThe real-time telemetry pipeline requires runtime binding. Ensure this window resides in a Chrome extension popup configured with permission parameters.โ, copyTextCommand = โOLLAMA_ORIGINS="" ollama serveโ, copyTagTxt = โMV3โ, copyHeaderTxt = โManifest Interface Schemaโ, copiedButtonTxt = โCopied Configurationโ, copyButtonTxt = โCopy Permission Manifestโ, }) => { const [copyState, setCopyState] = useState(false); const handleCopyManifest = () => { navigator.clipboard.writeText(copyTextCommand); setCopyState(true); setTimeout(() => setCopyState(false), 2000); }; const copyButtonTxtNode = copyState ? copiedButtonTxt : copyButtonTxt; return (
text-[9px] font-mono px-2 py-1 rounded-md transition-all bg-[#8B5CF6] text-white}> {copyTagTxt} <motion.button onClick={handleCopyManifest} className=โw-full py-2 bg-white/[0.06] hover:bg-white/[0.1] border border-white/10 text-xs font-mono font-medium rounded-xl flex items-center justify-center space-x-2 text-whiteโ> {copyButtonTxtNode} </motion.button> <ErrorUI> - Props: - headerDescTxt (type: any) [optional] - copyTextCommand (type: any) [optional] - copyTagTxt (type: any) [optional] - copyHeaderTxt (type: any) [optional] - copiedButtonTxt (type: any) [optional] - copyButtonTxt (type: any) [optional] - State Management: - Manages state copyState via setter setCopyState. - Hooks: Uses useState (Total Side-Effects: 0). - Rendered JSX Tree: <div>, <span>, <h1>, <p>, <button>, <motion.button> const ErrorUI = ({ headerDescTxt = โโฆโ, copyTextCommand = โโฆโ, copyTagTxt = โMV3โ, copyHeaderTxt = โโฆโ, copiedButtonTxt = โโฆโ, copyButtonTxt = โโฆโ })=>{ const handleCopyManifest = ()=>{ setTimeout(()=>setCopyState(false), 2000); โ/* โCopies the manifesto text to the clipboard and sets a temporary success state for two seconds.โ /โ; }; return ( text-[9px] font-mono px-2 py-1 rounded-md transition-all bg-[#8B5CF6] text-white}> {copyTagTxt} <motion.button onClick={handleCopyManifest}> {copyButtonTxtNode} </motion.button> text-[9px] ... {...} were left untouched. Static noise eliminated: Over 200 characters of utility noise were converted into short structural summaries. ๐ Benchmarks: Cold Runs vs. Warm Cache Runs We benchmarked urai-ecma on an Apple Silicon machine across different workloads using OpenAIโs native o200k_base BPE tokenizer: Benchmark 1: Real-World Monorepo Monolith (Chrome Agentic AI Extension) Metric Raw Project (TS/TSX) urai-ecma Output Total Reduction Token Volume (o200k_base) 209,757 tokens 36,153 tokens -82.76% ๐ Downstream LLM Context Exceeds local 64k limits Fits easily in local Ollama Usable on 8GB VRAM KV-Cache TTFT (Time-to-First-Token) ~18.4 seconds ~1.9 seconds ~9.6x Faster Benchmark 2: Tactical Project (plugin-api-docgen) โ Cold vs. Warm Performance Here are the terminal runs comparing an initial cold run (querying local Ollama) against subsequent warm runs (hitting the foyer hybrid cache): # RUN 1: Cold Execution (AST Parsing + Local Ollama Inference) time urai-ecma ๐ [urai-ecma] Starting AST Analysis on project: ./src ๐ Found 6 source file(s) for analysis. โ
[urai-ecma] Prompt successfully generated at: ./output.md ๐ [urai-ecma] Estimated Tokens in ./output.md: 1317 tokens ============================================================ ๐ TOKEN SAVINGS & OPTIMIZATION REPORT ============================================================ ๐ Raw Source Code (All JS/TS): 3023 tokens โก Optimized Output (output.md): 1317 tokens ------------------------------------------------------------ ๐ Reduction: -56.43% tokens saved! (Saved ~1706 tokens) ============================================================ urai-ecma 0.06s user 0.04s system 0% cpu 21.541 total # RUN 2: Warm Execution (AST Parsing + Foyer Hybrid Cache Hits) time urai-ecma ๐ [urai-ecma] Starting AST Analysis on project: ./src ๐ Found 6 source file(s) for analysis. โ
[urai-ecma] Prompt successfully generated at: ./output.md ๐ [urai-ecma] Estimated Tokens in ./output.md: 1299 tokens ============================================================ ๐ TOKEN SAVINGS & OPTIMIZATION REPORT ============================================================ ๐ Raw Source Code (All JS/TS): 3023 tokens โก Optimized Output (output.md): 1299 tokens ------------------------------------------------------------ ๐ Reduction: -57.03% tokens saved! (Saved ~1724 tokens) ============================================================ urai-ecma 0.03s user 0.01s system 92% cpu 0.049 total Performance Breakdown Execution Latency Comparison (plugin-api-docgen) โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Cold Run (Ollama Local Inference): โโโโโโโโโโโโโโโโโโโโ 21.541s Warm Run (Foyer Zstd Cache): โ 0.049s (49ms) โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Speedup Factor: ~439x Faster on Cache Hit! Cold Run (21.5s): The tool parses the AST in under 15ms, but waits on the local Ollama daemon to infer function docstrings sequentially. Warm Run (49ms - 68ms): On subsequent runs, Sha512_256 keys hit memory/disk cache entries with sub-millisecond latency. The entire repository is analyzed and emitted in less than 50 milliseconds! ๐ ๏ธ Tactical Comparison: urai-ecma vs. Existing Tools Feature repomix / gitingest code2prompt urai-ecma Engine Node.js / Python Rust Rust (SWC + Rayon) Parsing Strategy Naive String Concatenation Handlebars Templates True AST Traversal Tailwind Handling Preserves all noise (0% saved) Preserves all noise (0% saved) 4-Mode AST Pruning Structural Stubbing โ No โ No โ
Yes (is_structural_stub) API Route Tables โ No โ No โ
Auto-extracted (Next/Nest/Express) Component Analysis โ No โ No โ
Props, State, Hooks, JSX tree Privacy / Offline Dependent on API Offline string copy 100% Offline (Local Ollama/JSDoc) Token Savings 0% (Expands token size) 0% Up to 82.7% reduction โ ๏ธ Brutal Honesty: Trade-offs & When NOT to Use It No tool is a silver bullet. Because urai-ecma prunes function internals into architectural stubs, you need to understand when to use it and when to skip it: โ When NOT to Use urai-ecma Algorithmic Debugging: If you are trying to find an off-by-one bug in a sorting algorithm or a matrix multiplication loop, the function body is pruned. An LLM cannot debug math it cannot see! Memory Leak Profiling: If you need an LLM to inspect improper closure references or un-cleaned subscriptions inside an imperative block, you need raw source files. Writing Low-Level Unit Tests: If your unit tests need to mock intermediate local variables inside a 50-line imperative function, you need the full implementation. โ
When urai-ecma Shines Architectural Audits & Refactoring: High-level reviews of Next.js, React, or Express architectures. Autonomous AI Coding Agents (Claude Code, Aider, Cursor): Keeping an agentโs scratchpad clean so it navigates files without blowing through token limits. Spec-Driven Scaffolding: Passing an entire repositoryโs contracts and interfaces to an LLM to generate fresh implementations. Technical Interview Preparation: Asking questions about complex repositories without hallucinated paths or interfaces. Cross-Language Migrations (TypeScript to Rust/Go): Feeding the pure functional contract into an LLM without โJavaScript-ismsโ leaking into the generated target language. ๐ Installation & Quick Start urai-ecma is distributed as a single static binary with zero runtime dependencies. ๐ง macOS & Linux (Shell) curl โproto โ=httpsโ โtlsv1.2 -LsSf https://github.com/sanjaiyan-dev/urai-ecma/releases/download/v0.1.1/urai-ecma-installer.sh | sh ๐ช Windows (PowerShell) powershell -ExecutionPolicy Bypass -c โirm https://github.com/sanjaiyan-dev/urai-ecma/releases/download/v0.1.1/urai-ecma-installer.ps1 | iexโ ๐ฆ Node.js (Global Package Managers) npm install -g urai-ecma # or: pnpm add -g urai-ecma | bun add -g urai-ecma | deno add -g npm:urai-ecma ๐ฆ Rust (Cargo) cargo install urai-ecma โ๏ธ Configuration & SchemaStore Support Initialize a documented configuration file in your project root: urai-ecma create This creates urai.config.jsonc. Because urai-ecma is registered globally with SchemaStore, you get instant auto-complete and documentation in VS Code, WebStorm, IntelliJ, and Visual Studio: { โschemaโ: โhttps://www.schemastore.org/urai-ecma.jsonโ, // Source directory or file to analyze โinput_projectโ: โ./srcโ, // Target markdown prompt output โoutput_fileโ: โ./output.mdโ, // Local Ollama instance (Optional) โollama_endpointโ: โhttp://localhost:11434โ, โollama_modelnameโ: โgemma4โ, // Tailwind CSS mode: โremoveโ | โremove_aggrโ | โsummarizeโ | โpreserveโ โtailwind_modeโ: โremoveโ, โtailwind_thresholdโ: 96, // Summarize function bodies via JSDoc or Ollama โsummarize_functionsโ: true, โsummarize_functions_thresholdโ: 5, // Extract Express / Fastify / Next.js / NestJS routes โgenerate_route_tableโ: true, // React component introspection (Props, State, Hooks) โanalyze_react_componentsโ: true, // Generate ASCII tree & Mermaid ESM dependency graph โgenerate_file_graphโ: true } Run analysis anywhere: # Run with configuration file urai-ecma # Or run ad-hoc via CLI flags urai-ecma -i ./src -o prompt.md โtailwind-mode remove ๐ Try It Out & Contribute ๐ Documentation: https://sanjaiyan-dev.github.io/urai-ecma ๐ GitHub Repository: https://github.com/sanjaiyan-dev/urai-ecma ๐ค For AI Agents (llms.txt): https://sanjaiyan-dev.github.io/urai-ecma/llms-full.txt sanjaiyan-dev / urai-ecma AST-aware JS/TS codebase-to-prompt compiler in Rust. Uses SWC & Rayon to prune Tailwind bloat, retain structural stubs, extract Next.js/Nest routes, and compress repos into hyper-dense prompts for GPT-4o, Claude 3.5 & Ollama. Benchmarked with Tiktoken o200k for 80%+ token savings. ๐๏ธ URAI (เฎเฎฐเฏ) AST-Aware JS/TS Codebase-to-Prompt Engine for LLMs Transform bloated JavaScript & TypeScript repositories into hyper-dense, token-optimized context prompts. ๐๏ธ The Name Inspiration: The Art of โเฎเฎฐเฏ เฎเฎดเฏเฎคเฏเฎคเฎฒเฏโ (Urai Ezhuthudhal) In classical Tamil literary heritage, monumental epics and ancient treatisesโsuch as the Thirukkuแนaแธท, Tolkฤppiyam, and Cilappatikฤramโspan vast volumes of dense, poetic, and complex thought. To make these monumental texts intelligible without losing their depth, classical scholars (Uraiyฤsiriyars) practiced เฎเฎฐเฏ เฎเฎดเฏเฎคเฏเฎคเฎฒเฏ (Urai Ezhuthudhal): the disciplined art of writing a lucid, structured, and insightful commentary that distills the core essence, syntax, and architectural meaning of vast literature. The Modern Parallel Today, enterprise JavaScript and TypeScript codebases are the epic literatures of modern software. Spanning thousands of files across Next.js, React, Node.js, and TypeScript, they are laden with boilerplate, repetitive utility classes, and nested syntax. When feeding these systems to Large Language Models: โฆ View on GitHub If youโre sick of burning through API credits and watching your AI coding agents drown in static CSS strings, give urai-ecma a spin on your project. Drop your before-and-after token savings in the comments below!