Skip to content
All projects

AgentDesk

A full-stack AI assistant that calls real backend endpoints through LLM tool calling. It streams each step to the UI and asks before it changes any data.

Type
Side project
Tech stack
  • NestJS
  • React
  • PostgreSQL
  • OpenAI
  • Gemini
  • SSE
Illustration of AgentDesk

Why I built it

Most AI chat demos stop at answering questions. In a real business tool, the useful part is doing things: looking up an order, opening a ticket, pulling a number. I wanted to find out what it takes to let a model call real endpoints without it changing data it shouldn't.

How it works

You ask in plain language, and the agent picks the tools it needs, chaining several when a task has more than one step:

"Which orders from last week are still unpaid? Open a support ticket for the biggest one."

Here it runs search_orders, then create_ticket, and answers with the result. Each tool call shows up in the chat with its input and output, so you can follow exactly what happened.

  1. React UISSE stream
  2. NestJS agent serviceagent loop
  3. LLMOpenAI / Gemini
  4. Tool registryZod validation
  5. PostgreSQLPrisma
The agent sends messages and tool schemas to the LLM, validates each tool call with Zod, executes it and streams every step back to the UI.

The model can be OpenAI or Google Gemini, switched with one environment variable. Tokens and tool steps stream to the UI over Server-Sent Events, and conversations are stored per user in PostgreSQL. Login comes from my NestJS Production Starter.

The demo runs on sample data from a fictional shop, with tools like search_orders, get_order, get_customer, create_ticket, update_ticket and get_stats. Each tool is defined once with a Zod schema. That one schema validates the model's arguments at runtime and also produces the JSON schema the model sees, so the two can't drift apart. A new tool is about 20 lines:

export const getCustomer = defineTool({
  name: "get_customer",
  description: "Look up a customer by email or ID",
  schema: z.object({
    email: z.string().email().optional(),
    id: z.string().optional(),
  }),
  handler: async (args, ctx) => ctx.customers.find(args),
});

Keeping it in check

Tools that change data, such as create_ticket or refund_order, wait for the user to confirm. A model can misread what someone meant, and asking first is cheaper than undoing a refund. On top of that, each run has a cap on tool steps, and there is input validation, per-user rate limits and permissions per tool. Every run logs its prompts, tool calls, latency, token usage and cost.

I didn't use an agent framework. The loop is about 150 lines of TypeScript, which keeps it easy to debug and to explain. It's tested against a fake LLM that returns scripted tool calls, so the tests are repeatable and cost nothing.

What surprised me

The loop itself turned out to be the small part. Most of the work went into validation, error handling and limits. I also found that clear tool names and schema fields improved the model's tool choice more than longer system instructions did.