All resources
Daily Learning · Ad-hoc Posts

From calling a model to managing an interaction

Tuesday, 30 June 2026

🎯 You'll understand how AI apps are moving from "send a prompt, get a reply" to a full managed conversation — and the big design choice that comes with it: how much of that work to hand to the AI provider versus keep in your own backend.

Let's learn how building with AI is changing — slowly but deeply. The model itself isn't the interesting part today. The interesting part is everything around the model.

Let me build this up one piece at a time.


Step 1: The simplest way to use an AI model

In the beginning, using an AI model is wonderfully simple. You send it some text. It sends some text back.

You ───"What's a good name for a cat?"───► Model
You ◄──────"How about Whiskers?"────────── Model

That single round-trip is called a prompt.

What is a prompt?

A prompt is the text you send to a model to get a response. Example: "Summarize this email in one sentence."

For a long time, the prompt was the whole game. You wrote a good prompt, you got a good answer. People even called it "prompt engineering."

But real apps need a lot more than one question and one answer.


Step 2: Real apps need a conversation, not one message

Imagine a customer support assistant. The customer says "I want a refund." The assistant needs to:

  • Remember what was said earlier
  • Look up the customer's order in your database
  • Maybe charge or refund money
  • Keep the conversation flowing smoothly

One prompt can't do that. You need a loop — a back-and-forth that keeps going until the job is done.

This loop is often called the agent loop.

What is an agent loop?

An agent loop is the repeating cycle where the model thinks, asks for help (tools), gets results, and continues — until it has a final answer. Think of a chef who keeps cooking, tasting, and adjusting until the dish is ready.

Here's what that loop looks like by hand:

1. Send the user's message to the model
2. Read the model's reply
3. Does it want to use a tool? 
      → Yes: run the tool, send the result back, go to step 2
      → No:  show the final answer to the user

Step 3: The "glue code" — where the real work hides

Here's the surprising truth: that loop above is small to describe but huge to build well.

Every single step needs careful code:

  • Call the model — and handle timeouts and errors.
  • Parse the response — figure out: is this a final answer, or a request to use a tool?
  • Run the tool — call your database or payment system safely.
  • Feed the result back — re-package it so the model understands.
  • Stream the tokens — show words as they appear, not all at once.
  • Track session state — remember the whole conversation.
  • Log everything — so you can debug and audit later.

All of this code that sits between the model calls is what engineers call glue code.

What is glue code?

Glue code is the plumbing that connects pieces together. It's not the exciting part — it's the wiring that makes the exciting parts actually work. Like the pipes behind your kitchen wall: invisible, boring, and absolutely essential.

In production AI systems, most of the real complexity lives in the glue code — not in the model call itself.

Let me name the six big jobs this glue code does, because they matter for the next step:

SESSIONS          remembering the whole conversation
EVENTS            tracking each thing that happens
TOOL CALLS        letting the model use your functions
MULTIMODAL CONTEXT handling text, images, and more
STREAMING         sending words out as they're generated
PERSISTENCE       saving everything so it survives

Quick word: what is "multimodal"?

Multimodal means more than one type of input or output — text, images, audio, video — not just words. Example: sending the model a photo of a receipt and asking "what's the total?"

Quick word: what is "persistence"?

Persistence means saving data so it doesn't vanish when the program stops. Like writing in a notebook instead of just remembering in your head.


Step 4: The big shift — the platform takes the loop

Now here's the change worth understanding.

Google shipped the Gemini Interactions API. The key idea isn't a smarter model. It's that the loop itself — all that glue code — moves inside the API.

Instead of you writing:

"Call a model and get a response."

You now write:

"Manage an interaction."

What is an "interaction" here?

An interaction is the whole managed exchange — the full back-and-forth, with memory, tools, streaming, and saving — treated as one thing the platform handles for you. Compare it to a prompt, which is just one message.

Picture two ways of running a restaurant:

OLD WAY: You're the waiter, the cook, the cashier,
         AND you wash the dishes. You run between
         every station yourself.

NEW WAY: You place the order. The kitchen handles
         cooking, plating, timing, and cleanup.
         You just manage the customer relationship.

So those six jobs — sessions, events, tool calls, multimodal context, streaming, persistence — stop being your code. They become first-class parts of the API.

What does "first-class" mean?

A feature is first-class when the system treats it as a built-in, official thing — with proper support — rather than something you bolt on yourself. Example: a car with built-in GPS (first-class) vs. taping your phone to the dashboard (bolted on).


Step 5: What the new flow actually looks like

Let me walk through one full interaction, step by step, the way the Interactions API runs it.

┌─────────────────────────────────────────────────┐
│ 1. Your app sends:                               │
│      user input + session context                │
│                  │                               │
│                  ▼                               │
│ 2. The model decides:                            │
│      "Answer directly"  OR  "I need a tool"      │
│                  │                               │
│         ┌────────┴────────┐                      │
│         ▼                 ▼                      │
│   Direct answer    "Run the refund tool"         │
│                          │                       │
│                          ▼                       │
│ 3. YOUR backend runs the tool                    │
│      under your auth, data rules, guardrails     │
│                          │                       │
│                          ▼                       │
│ 4. Result flows back into the interaction        │
│                          │                       │
│                          ▼                       │
│ 5. Model produces the next event or final answer │
│                                                  │
│   ── The whole thing streams, is observable,     │
│      and persists ──                             │
└─────────────────────────────────────────────────┘

Notice something important in step 3: even though the platform manages the conversation, your backend still runs the actual tools.

What is a "tool call"?

A tool call is when the model says "I can't do this myself — please run this function for me." Example: the model asks your code to run getOrder(orderId) because only your system can reach your database.

What is "auth"?

Auth (short for authorization/authentication) is checking who is allowed to do what. Example: making sure this user is actually allowed to refund this order.

What are "guardrails"?

Guardrails are the safety rules that stop the system doing something harmful or wrong. Example: "Never refund more than the original payment."

What does "observable" mean?

Observable means you can see what's happening inside — every step, every event — so you can debug and audit. Like a glass-walled kitchen where you can watch every dish being made.


Step 6: The real win — and the real catch

The win is clear: less hand-rolled orchestration.

What is "orchestration"?

Orchestration is coordinating many moving parts so they work together in the right order. Like a conductor keeping every musician in time.

Before, you orchestrated the whole loop yourself. Now the platform does much of it. That's less code to write, fewer bugs, faster building.

But here comes the catch — and this is the part a thoughtful builder must sit with.

Every capability you push into their API is glue code you delete today — and a dependency you inherit tomorrow.

What is a "dependency"?

A dependency is something your system relies on that you don't control. Example: if your app only works because of one company's API, you depend on them — their prices, their rules, their uptime.

So now you face a genuine fork in the road. Two forces pull in opposite directions:

CONVENIENCE  ◄───────────────►  CONTROL

Let the provider          Keep it in your backend
handle it:                yourself:
+ less code               + you own your state
+ faster to ship          + you own security
+ fewer bugs to write     + you handle failures your way
                          + you can switch providers

Step 7: The architect's question

This is the decision that defines modern AI systems. For each of those six capabilities, you must ask:

Should this live inside the provider's API, or stay in my own backend?

Things you might happily hand over:

  • Streaming the tokens
  • Tracking the basic event list

Things you might insist on keeping yourself:

  • State — the source of truth about your data
  • Security — who can do what, under your rules
  • Failure handling — what happens when something breaks
  • Provider freedom — the ability to switch to a different model company later

There's no single right answer. A small startup might hand over almost everything for speed. A bank might keep nearly everything for control.

The model used to be the product. Increasingly, the interaction boundary is.

The interaction boundary is simply the line you draw: what runs in their API versus what runs in yours. Where you put that line is now one of the most important design decisions you'll make.


The core lessons

  • A prompt is one message; an interaction is the whole managed conversation — with memory, tools, streaming, and saving.
  • The hard part of AI apps was never the model call — it was the glue code in the agent loop: sessions, events, tool calls, multimodal context, streaming, and persistence.
  • The Gemini Interactions API moves that loop into the platform, making those six jobs first-class features instead of code you write by hand.
  • Your backend still runs the real tools — under your auth, data rules, and guardrails — so sensitive work stays under your control.
  • Convenience and control pull in opposite directions. Handing work to the provider deletes glue code today but creates a dependency tomorrow.
  • The key decision is where you draw the interaction boundary — what lives in their API, and what you keep in your own backend for state, security, failure handling, and the freedom to switch providers.

Want this kind of thinking applied to your business?

Book a free 30-minute discovery call — we’ll show you your highest-value first automation, no jargon, no obligation.

Book a discovery call
From calling a model to managing an interaction | Kingsmen Daily Learning