You do not need Photoshop or a subscription to put a product on a clean white background. A laptop with an NVIDIA graphics card, ComfyUI and one open model do it in a few seconds per photo, and a short graph handles the rest: the crop, the margin and the proportions. I built this for the shop of my garage, I use it for new products too, and it all runs on my own laptop.
Why I stopped using online tools
I need packshots often, and not for one project only. The parts in my garage come from several donor cars, and they get photographed wherever they happen to lie: in a box, on the lawn, on the workshop floor. New products need the same treatment. A shop, a catalogue or an offer looks better when every cover is the product alone on white.
I tried online background removers first, among them iloveimg and Adobe Firefly’s remove background. The second one did not work well for me. The best result still comes from selecting the part by hand in Photoshop with the Pen Tool. That is the right method for a professional photo session, and the wrong one for hundreds of product thumbnails: it takes too long.
So I did it locally. ComfyUI turned out to be a good fit, because the whole job is a graph that I build once and then run on every photo.
I describe the shop itself, with the review panel and the warehouse behind it, in the garage inventory system case study; here I stay with the packshots.
What ComfyUI is
ComfyUI is a free, open-source program that runs AI image models on your own computer. Instead of a single “remove background” button you build a graph: you connect nodes, each doing one thing (load a photo, cut out the background, resize), and the whole graph is your tool. It runs as a local server with a browser interface and an HTTP API, so the same graph works by hand and from a script.
What you need
| What I use | |
|---|---|
| Laptop | Lenovo, Intel Core Ultra 9 275HX, 64 GB RAM, Windows 11 |
| Graphics card | NVIDIA RTX PRO 3000 Blackwell Laptop GPU, 12 GB VRAM |
| ComfyUI | 0.13.0 (Python 3.13, PyTorch 2.10 with CUDA) |
| Model | BiRefNet-General |
| Custom nodes | ComfyUI_LayerStyle_Advance (runs BiRefNet), ComfyUI_essentials (MaskBoundingBox+) |
I have not tried it on weaker hardware, so I will not promise numbers there.
How the graph works

The numbers below match the labels on the screenshot.
- Source photo. A
LoadImagenode. It is the only place the graph changes from one photo to the next. - BiRefNet model. Loads
BiRefNet-General. - Cut out the part.
BiRefNetUltraV2returns a mask: white where the part is, black elsewhere. I useGuidedFilterfor the edge, erode 4, dilate 2, black point 0.01 and white point 0.99. The node works on at most 4 megapixels internally. - White canvas, then the part on white (4 and 5). An empty white canvas the size of the photo, and the original pasted onto it through the mask, so everything outside the mask stays white.
- Crop to the outline (6).
MaskBoundingBox+cuts the image to the box around the mask. This is why the margin is even, whatever the framing of the original. - Fit it into the final frame (7a–7c and 8a–8c).
ResizeAndPadImagescales the part into 1800×1300, and a composite node places it at x=100, y=100 on a white 2000×1500 canvas, which leaves a 100 px margin on every side. A second branch does the same into 2000×2000.
I make two files from one pass because in my shop the tiles are 3:2 and the gallery is 4:3. A 4:3 image fits both without cutting the part, and the square is for square thumbnails. The two PreviewImage nodes are titled packshot 4:3 and packshot 1:1. Keep those titles if you call the graph from a script, because that is how the script finds the outputs.
Download the graph: packshot-birefnet.workflow.json opens in ComfyUI (drop it on the canvas), and packshot-birefnet.api.json is the same graph in the API format for scripts.
A transparent background instead of white
The same mask gives you a transparent background. Invert it with InvertMask, feed it to JoinImageWithAlpha together with the original photo (that node inverts the mask itself, so the double inversion keeps the part opaque) and save the result as PNG. The cover of this article in dark mode is made this way. Only the cropping and centring I did outside ComfyUI, with ImageMagick.
Pitfalls I hit
The BiRefNet-General weights have to be downloaded by hand. The node fetches only a different model on its own:
python -c "from huggingface_hub import snapshot_download; snapshot_download(repo_id='ZhengPeng7/BiRefNet', local_dir=r'<ComfyUI>\models\BiRefNet\BiRefNet-General', ignore_patterns=['*.md', '*.txt'])"With transformers 5 the BiRefNet node failed for me with
Input type (float) and bias type (c10::Half) should be the same. Passingdtype=torch.float32tofrom_pretrained(...)in the node’sbirefnet_ultra_v2.pyfixed it. A newer version of the node may not need this, and an update can undo the change.Do not crop with
CropByMask V2. Inmin_bounding_rectmode it keeps only the largest contour, and it cut off the headrests of car seats that are joined to the backrest by thin rods.
What it looks like while it runs

You press Run and watch. The progress bar fills at the top of the window, the node being computed is outlined in blue, and the job queue on the right lists every finished job with its time. The tab title shows the percentage and the name of the current node, which is handy when ComfyUI sits in the background.
On my laptop a photo of about 2000 px, the size I usually keep, takes 2–4 seconds. A full 3800 px phone photo takes about 5 seconds. The first run after starting ComfyUI is slower, about 6 seconds, because the model has to load first.
Examples
All of these are the raw output of the graph, with no retouching afterwards. The source photo is on the left.


This one came out especially well. The Audi engine cover lay on packing paper on a patterned, reflective surface, with shipping labels next to it, and only the cover was cut out, with the rings and the TFSI lettering sharp. Here you see the square output; the 4:3 one comes from the same pass.



The model is not thrown by a grey table, the grid of a cutting mat, shadows on the grass or a red clamp either. The last pair shows where it goes wrong:

The filter stands on its box, and the box stayed in the picture, because the model does not know which object is for sale.
From one photo to a few hundred
You can call the graph from any language, because ComfyUI exposes an HTTP API. A script has to do three things: upload the photo and point the LoadImage node at it, queue the graph and wait for it to finish, then download the two outputs. This is the whole thing in Python, with the API version of the graph saved next to it:
import json
import sys
import time
from pathlib import Path
import requests
COMFY = "http://127.0.0.1:8188"
WORKFLOW = Path("birefnet-2000-white.api.json")
def packshot(photo: Path, out_dir: Path) -> None:
graph = json.loads(WORKFLOW.read_text())
# 1. Upload the photo and point the LoadImage node at it.
with photo.open("rb") as f:
name = requests.post(f"{COMFY}/upload/image", files={"image": f}, data={"overwrite": "true"}).json()["name"]
load = next(k for k, n in graph.items() if n["class_type"] == "LoadImage")
graph[load]["inputs"]["image"] = name
# 2. Queue the graph and wait for it to finish.
prompt_id = requests.post(f"{COMFY}/prompt", json={"prompt": graph}).json()["prompt_id"]
while not (entry := requests.get(f"{COMFY}/history/{prompt_id}").json().get(prompt_id, {})).get("status", {}).get("completed"):
time.sleep(0.5)
# 3. Save both outputs, found by the titles given to the PreviewImage nodes.
for node_id, node in graph.items():
if node["class_type"] == "PreviewImage":
image = entry["outputs"][node_id]["images"][0]
data = requests.get(f"{COMFY}/view", params=image).content
suffix = node["_meta"]["title"].split()[-1].replace(":", "x") # "4x3" or "1x1"
(out_dir / f"{photo.stem}-{suffix}.png").write_bytes(data)
if __name__ == "__main__":
out = Path("packshots")
out.mkdir(exist_ok=True)
for path in sys.argv[1:]:
started = time.time()
packshot(Path(path), out)
print(f"{path}: {time.time() - started:.1f} s")Download packshot.py.txt, rename it to packshot.py, put the API graph next to it as birefnet-2000-white.api.json and run python packshot.py photos/*.jpg. On my laptop it prints 2–3 seconds per photo and writes name-4x3.png and name-1x1.png into a packshots folder.
When the site is on a server without a GPU
The garage shop runs on a server with no graphics card, so I did not move the work there. The laptop asks the server for waiting jobs instead, one at a time, and the server never connects to the laptop. In my Laravel code the loop looks like this (abridged):
$job = $api->claim(1)[0] ?? null; // ask production for a queued render
$packshot = $renderer->render($api->download($job['source_url'])); // run the graph in ComfyUI
$upload = $api->uploadUrl($productId); // presigned URL on the CDN
$api->upload($upload['upload_url'], $packshot->jpeg, $upload['content_type']);
$api->upload($upload['square_upload_url'], $packshot->squareJpeg, $upload['content_type']);
$api->complete($productId, $upload['key'], $packshot->workflowVersion->value);If the laptop is off, the jobs wait in the queue and nothing is lost. The graph file is versioned in the repository, and every packshot records which version of the graph made it, so after a change I know which photos to redo.
Where the model gets it wrong, so I check every result
The model removes everything that looks like foreground. In practice that means a few kinds of photos give a bad result:
- A part marked with a frame or an arrow on a larger assembly. An intercooler circled on a stack of radiators comes out as the whole stack.
- A part photographed on a car. Wheels on a car come out as the whole car.
- Other objects or labels in the frame. They are cut out together with the part, like the cardboard box above.
- Openings and holes. The model decides on its own whether what shows through an opening is part of the object. The engine cover above has a round hole: on the 1200×1600 copy the graph cut it out, on the 3028×4038 original it left the packing foil visible in it. The result also changed with the size I scaled the photo to, so check the holes and, if one matters, draw it by hand.
- Clear lenses, thin wires and chrome. These can be misjudged, so no packshot is published without my approval.
A finished packshot waits in the admin panel next to its source photo and becomes the cover only after I approve it. If it is wrong, I order it again from another photo where the part lies alone, or I reject it and the listing keeps the original photo.
Comparison with online tools
To test this on something difficult, I ran the same photos through two online tools, iloveimg (free version, no account) and Adobe Firefly’s remove background (also without logging in), and through my graph. Both photos went through all three tools.
An engine cover on packing paper
The first photo is the Audi engine cover from above. It lies on packing paper on a patterned, reflective metal surface, with shipping labels next to it.
Adobe Firefly took away only the outer edge of the scene and kept 51% of the frame. The packing paper, the patterned surface and the labels stayed, so the result is still a photo of the whole scene with the cover somewhere in it. iloveimg and the ComfyUI graph both isolated the cover itself, keeping about 15% of the frame. Adobe got the original file (3028×4038), iloveimg and the graph a copy scaled down to 1200×1600.

An air filter housing on grass
The second photo is an air filter housing lying on grass in strong sun, with a hard shadow, a thin pipe and a bundle of wires.

All three tools handled the overall shape, including the pipe and the wires. iloveimg and Adobe return the cutout in the original frame, so the part stays in it and still needs cropping and centring. The graph does that in the same pass and returns a finished 4:3 file with an even margin.

The cutouts of iloveimg and ComfyUI overlap in 99.5% of their pixels, and Adobe’s overlaps with each of them in about 98%. Zoomed in, the differences are about sharpness: iloveimg’s edges are a little softer than the graph’s, and Adobe’s are the softest, with bits of grass left between the wires. Adobe also returned the image at 800×600, half the size of the photo I gave it. For the cutting alone iloveimg and the graph are level. What the graph adds is the finished framing and the automation, not a better cutout, and what separates it from iloveimg is the free plan: the service itself says that unlimited image processing comes only with Premium. Adobe Firefly also failed on the first photo, but two photos and one attempt each are an observation, not a ranking.
When something else is the better choice
- Photoshop with the Pen Tool still gives the best edge. It is the right choice for a professional photo session, where a single picture matters and there is time to perfect it.
- An online remover is quicker for a single photo, because there is nothing to install. With many photos you have to reckon with the limits of the free plan or with a paid one.
- This graph pays off when you do it often and have an NVIDIA card: the setup takes some time once, and every photo after that takes a few seconds.
Summary
- ComfyUI with BiRefNet turns a photo of a part into a packshot on white in 2–4 seconds on a laptop GPU. It is free, and the processing happens on your own computer.
- The graph is short: cut out, paste on white, crop to the outline, fit into a frame with a margin. One pass gives a 4:3 file and a square one.
- A short Python script runs it on a whole folder, and a small worker can feed it from a shop or site on a server without a GPU.
- The model does not know which object is for sale, so every result needs a quick look before it goes live.
- For a single photo or for a professional session, an online tool or Photoshop is the better tool.
I use the same tool to upscale photos, too. How do you make your packshots: by hand, with an online tool, or locally? I am curious what works for you and what this graph is missing.




