Wire It, Run It, Deploy It: AI Workflows in Gradio
2026-08-26 · Hugging Face
Overview
Most interesting AI apps are pipelines. You generate an image, then remove its background or edit it into something new. You write a script, then generate a voice for it, or swap the voice while keeping the script. We usually wire these steps together in Python, and when something looks off we resort to print-debugging to find the offending step.
gr.Workflow, built into Gradio, makes the pipeline the interface. You describe your steps as a graph of typed nodes, and Gradio serves a drag-and-drop canvas where every node is runnable and every intermediate result is visible. The same graph is also a REST API and deploys to Hugging Face Spaces with a single command.
Examples in Action
Each app below is a live Hugging Face Space you can open, run, and duplicate.
Edit an Image
Upload an image, type an edit ("turn it into a snowy winter scene", "add sunglasses", "make the car red"), and get the edited photo back. The whole app is a single node calling Qwen-Image-Edit on Hugging Face Inference Providers.
Chain Real Models into a Media Studio
One graph, three pipelines:
- Start with a prompt, generate an image with FLUX, then pass it to a background-removal Space to make a sticker.
- A topic becomes a voiceover through a text-to-speech Space.
- The same topic becomes a catchy episode title through an LLM call.
That's one canvas, two model calls via Inference Providers, and two Space calls. Since it's a workflow, each output also gets its own REST endpoint: `/sticker`, `/voiceover`, and `/episode_title`, callable directly from code.
Fan-out Image Generation in Parallel
Type one idea and it becomes a set of artwork at once: a base image from FLUX, two AI re-imaginings (a soft watercolor and a neon cyberpunk take), and an LLM-written gallery title. This is the fan-out pattern: one idea feeds multiple operators simultaneously, all generating in parallel.
Profile a Hugging Face Dataset
Type a dataset ID such as `stanfordnlp/imdb`, and a single input fans out to four operator nodes that analyze the dataset live via the Datasets Server API: an overview card, a preview of the first rows, per-column statistics, and a distribution chart, all computed independently and in parallel.
Run Your Own GPU Model
An fn node is just Python, so it can run a model inside the Space on a GPU. Decorate the bound function with `@spaces.GPU` and ZeroGPU grabs a GPU for that call, runs the model, and releases it. A demo animates a still image using Lightricks/LTX-Video loaded through Diffusers, running entirely through one node.
How It Works
Every workflow is a graph with three kinds of nodes:
- References: your inputs.
- Operators: the steps that do work, whether your own Python function, a model on Inference Providers, another Gradio Space, or a row from a Hub dataset.
- Subjects: your outputs.
You connect them by dragging between typed ports, hit Run, and watch each result appear in place.
Call It from Code
Every workflow is also an API with no extra work. Each output becomes a REST endpoint named after its label:
from gradio_client import Client
client = Client("ysharma/gr-workflow-multi-endpoint-API")
print(client.predict("hello there friend", api_name="/word_count")) # -> 3
print(client.predict(20, api_name="/fahrenheit")) # -> 68.0
Endpoints that call a model or Space run under a Hugging Face token, so pass one when creating the client. Every endpoint is also reachable over curl via plain HTTP.
Build Your Own
The fastest way in is to open any demo, click Duplicate, and start rewiring. From Python it's short:
import gradio as gr
def your_function(text: str) -> str:
pass
gr.Workflow(bind=[your_function]).launch()
For the full walkthrough, operator kinds, JSON schema, and reusable patterns, see the official gr.Workflow guide in the Gradio docs. You can even build something as involved as AUTOMATIC1111 with gr.Workflow, which a future post will cover step by step.