PraxisAgent: Autonomous Open-AI Operations Worker Built for My Friend Sarah
This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend What I Built I built PraxisAgent for my friend Sarah, who handles billing and clinic operations. Every Friday afternoon, Sarah faces the same nightmare: dozens of vendor invoices scattered across messy folders, unorganized scanned PDFs with scrambled filenames (scan_0003.pdf), and outdated ERP / healthcare web portals that require tedious, error-prone manual data entry. She was constantly stressed about: Accidental duplicate payments on invoices that were already settled. Typos in payment amounts and dates when copying from scanned PDFs into web forms. Losing hours of her weekend to repetitive browser clicking and form-filling. PraxisAgent is an autonomous, open-source operations agent that takes plain-English operational requests (e.g., “Find Acme’s latest bill in the invoices folder and enter it into the ERP”), autonomously locates and parses unstructured documents (PDFs, JSONs, text tickets), drives a real browser session to populate the portal, and stops at an interactive Human-in-the-Loop policy gate before submitting irreversible actions. Bonus — Handing it over to Sarah: When I showed Sarah the agent running locally on her workstation and asked her to test it on a batch of Acme Corp invoices, she watched it parse the scanned PDF, filter out an already-paid duplicate, fill out the ERP form in Chromium, and pause right at the “Submit Voucher” screen asking for her approval. Her reaction: “Wait, so I never have to manually re-type 20-digit voucher IDs while squinting at scanned PDFs again? And it physically won’t send money unless I click approve? This just gave me my Friday evenings back.” Demo Here is a look at PraxisAgent in action: Key Capabilities in Action: Unstructured Multi-Format Extraction: Parses native PDFs (pdf-parse) and unstructured JSONs, ignoring misleading file modification timestamps and reading ground-truth dates directly inside the document. Numbered Element DOM Distillation: Compresses live browser pages into compact [ref] selectors (~150–300 tokens/step), allowing fast, reliable form navigation without flooding the LLM context window. Deterministic Policy Safety Gate: Intercepts state-mutating actions (submit, approve, pay), extracts form field values, and prompts the user for interactive confirmation before executing. Independent Out-of-Band State Verifier: An isolated background auditor queries the target portal’s database (/__state) directly to cryptographically prove that records were created with exact amounts and no duplicate submissions occurred. Code The entire codebase is open-source, fully typed in TypeScript, and runs with zero external agent frameworks: jaiswalabhishek377 / praxis-agent The autonomous agent that bridges intent and verified execution. PraxisAgent — Autonomous AI Operations Worker An autonomous enterprise operations agent that executes end-to-end IT, ERP, and healthcare workflows across real browser portals and local filesystems—featuring element-referenced DOM perception, code-level approval gates, multi-model failover, and independent state verification. ⚡ Core Highlights Zero-Framework ReAct Loop: Hand-written TypeScript state machine with strict Zod schema validation and 1-turn self-repair—zero LangChain/CrewAI dependencies. Ephemeral DOM Perception: Distills live pages into compact numbered element references ([ref] IDs, ~150–300 tokens/step). Prunes stale DOM trees each turn while preserving action history to prevent context degradation. Deterministic HITL Policy Gate: Code-level security firewall that intercepts irreversible actions (submit, approve, pay, confirm), displays extracted form values, and fails closed (deny) in non-interactive terminals. SHA-256 Loop Prevention: Computes SHA-256 hashes of (page_state, action, params) and automatically aborts after 3 duplicate actions without progress. Multi-Model Provider Cascade: 4-tier Gemini… View on GitHub How I Built It PraxisAgent is built from the ground up without heavy third-party agent wrappers (zero LangChain, zero CrewAI): Zero-Framework ReAct Loop: A lightweight, deterministic TypeScript execution engine with strict Zod schema validation and 1-turn automated JSON self-repair. Open-Weight AI at the Core: Powered by open-weight models including Llama 3.3 70B and Gemma 2 via local inference (Ollama / vLLM) and Groq high-speed inference. Built with a resilient Multi-Model Provider Cascade: automatically fails over across tiers when hitting rate limits or provider downtime, ensuring 100% workflow completion without crashing. Playwright Headless/Headed Browser Perception: Operates directly on live web portals via compact element-referenced snapshots, restricted strictly to sandboxed origin allowlists. SHA-256 State-Action Loop Prevention: Hashes (page_state, action, params) on each step to immediately abort if an agent gets caught in a loop. Auditable Artifact Dossier: Every run automatically saves an execution trace (runs/.jsonl), step-by-step screenshots before/after actions, and an independent database verification dossier (audit_dossier.json). Why Does Open Innovation Matter? For an enterprise operations and healthcare tool, open innovation isn’t just a nice-to-have—it is an absolute requirement: Data Sovereignty & Privacy (HIPAA & Financial Records): Invoices, vendor tax IDs, and patient medical claim records contain sensitive, confidential information. Closed proprietary APIs often retain customer data for retraining or inspect prompts in black-box cloud environments. With open-weight models (running locally or in self-hosted instances), sensitive financial and healthcare data never leaves the organization’s control. Zero Vendor Lock-in & Model Swappability: Closed models frequently change weights, undergo silent behavioral drift, or deprecate endpoints overnight. By building on an open-weight foundation, we can swap between open models (Llama, Gemma, or custom fine-tunes) seamlessly without rewriting our tool schemas or agent loop. Predictable Zero-Marginal-Cost Execution: Routine administrative tasks like invoice processing run thousands of times a month. Running inference locally or on dedicated open-weight instances eliminates unpredictable per-token SaaS subscription spikes. Auditability & Code-Level Safety: Because the agent runtime is 100% open-source TypeScript, our safety firewall (the Human-in-the-Loop policy gate) is enforced at the code layer, not as an afterthought prompt instruction. My Agent Session PraxisAgent was designed, tested, and benchmarked across 9 end-to-end chaos engineering scenarios (achieving a 9/9 verified pass rate across validation errors, session dropouts, and duplicate invoice filters). Full evaluation logs and architectural diagrams are available in eval-results.md and architecture.mmd. Prize Categories Best Use of Gemma: Can run open-weight Gemma models locally for air-gapped, zero-data-leakage document extraction and browser operation, locally ollama powered by gemma2-9b-it within its Multi-Model Provider Cascade to run fast, open-weights fallback inference for browser automation and document parsing. Best Use of Entire: PraxisAgent records step-by-step JSONL execution traces (runs/.jsonl) and database verification audit dossiers for complete operational transparency. Best Use of GitHub Copilot: Used Copilot to write code.