Black Forest Labs trained one model on images, video, and audio together, then taught the same network to move a robot arm. The crew is collapsing into the camera.
Issue No. 078Friday, July 24, 2026Running time ≈ 4 min
FeatureTonight's PictureTC 00:00:00
One model learns picture, sound, and motion
Black Forest Labs announced FLUX 3, a frontier model that learns images, video, and audio inside a single architecture rather than stitching three specialists together. It renders up to 20 seconds of video with native dialogue, effects, and score, holds characters across multi-shot sequences, and BFL calls out strong typography and multilingual text rendering, the parts designers stress test first. Video is in early access now; image generation and an open-weight FLUX 3 Dev follow in the coming weeks.
The stranger half of the launch is FLUX-mimic, built with mimic robotics: the same backbone decodes robot actions from its learned video representations, and it is already running manipulation tasks on Audi production lines after fine-tuning on roughly 30 minutes of task data. Early preference tests put FLUX 3 video ahead of Runway Gen-4.5 in 77 percent of head-to-head comparisons.
One negative, three exposures. Picture, sound, and motion come off the same roll.Today's Art Direction
The Screening Room
A motion studio's landing page shot at night: didone title cards over tungsten haze, timecoded reels, an end-credit colophon.
The archetype is a production studio's site, the page a freelancer ships to a director of photography or an animation shop. The register is cinema: a full-bleed photographic ground, film-title typography in a high-contrast didone, and section markers borrowed from the edit bay, reel numbers and timecodes instead of icons. Everything readable sits on a two-layer scrim, a duotone tint plus a directional gradient, so the photograph stays dark and alive under live text.
title card
two-layer scrim
duotone tint
filmography row
timecode marker
tungsten haze
end credits
Reel 01ToolingTC 00:01:20
Claude Opus 5 lands at Opus prices
Anthropic released Claude Opus 5, available today on every platform at the same $5 and $25 per million tokens as Opus 4.8. It comes close to Fable 5's frontier intelligence at half the price, posts state-of-the-art marks on Frontier-Bench and GDPval-AA, and becomes the new default on Claude Max.
Webflow retires the bridge
Webflow's MCP server hit 2.0: most element, component, style, and variable operations no longer need a live Designer session, and every action an agent takes runs through workspace roles, permissions, and audit logs with explicit AI attribution. Webflow says nearly 90 percent of its MCP users connect through Claude.
Vercel's MCP server deploys code now
Vercel MCP crossed the line from inspection to action: a connected agent can deploy code directly. Eve, the agent that took the on-call shift in issue 076, also picked up installable extensions this week, GitHub tools first on the shelf.
Reel 02TechniqueTC 00:02:10
Figma puts agents on the vulnerability beat
Figma's security team documented how it runs agents at three points: code generation, pull request review, and audits of the back catalog, all steered by one shared security policy. Nothing merges without human review, and the engineers moved up a level, from triaging single bugs to writing the policy that catches hundreds.
Foldkit pitches correctness as the feature
Foldkit is a new front-end framework built on Effect and architected like Elm: model, message, update, view, with effects tracked in the type system. Worth a look as a counterweight while agents write a growing share of the front end.
Reel 03WorkflowTC 00:03:05
Stack Overflow names the bottleneck: context
Stack Overflow's No Dumb Questions series takes on why capable models still stall inside real teams, and lands on context: the model is rarely the constraint, the information reaching it is. Context engineering gets treated as a discipline of its own, with practices a working team can adopt.
A 150MB Copilot binary froze FreeBSD ports
A maintainer accidentally committed the entire Linux binary of GitHub Copilot CLI to the FreeBSD ports tree, breaking the GitHub mirror's file limit and raising license questions. Core froze the whole tree for days while the rollback landed, a tidy parable about what one unreviewed artifact does to a shared commons.
Borrow this pattern
The two-layer scrim
This page runs its headline straight over a photograph, and the move that keeps it honest is deliberately boring. Layer one is a duotone tint: fill the whole image with the page's ground color at 30 to 40 percent opacity, so the photo adopts the palette instead of fighting it. Layer two is a directional gradient scrim, near 85 percent opacity behind the text block, falling to clear where the image should breathe.
Then test the smallest line of type against the busiest region of the photo at desktop and 390px widths; when it misses WCAG AA, darken the nearest gradient stop rather than the whole image. Use it any time a client hands you one good photograph and a headline that has to live on top of it.
BoothPrompt LabTC 00:03:40
One prompt returns the whole screening room, scrim recipe included.
Build a one-page landing site for a motion production studio, art directed as a night screening room. Ground the page in warm near-black (#17100C) with cream text (#F4EAD9), tungsten amber (#E3A93E) for links and marks, curtain crimson (#A93327) reserved for thin rules, and warm tan (#A08D7B) for quiet labels. Pair a high-contrast didone display face (Bodoni Moda) with a plain grotesque body (Archivo) at 19px and 1.7 line height, plus a monospace (Space Mono) for timecodes and captions.
Hero: a full-bleed photographic still of a dark film set with warm practical lights, behind a centered title card: a letterspaced monospace kicker, an oversized didone film title, a two-sentence logline, and a stacked three-line credit block for version, date, and running time. Never a boxed metadata grid. Keep the type readable with a two-layer scrim: tint the photo with the ground color at about 35 percent, then run a vertical gradient from 85 percent opacity behind the text down to clear, and verify WCAG AA over the busiest part of the image at desktop and 390px widths.
Run sections as numbered reels. Each opens with a ruled header row: reel number left in amber monospace, section name centered in letterspaced caps, timecode right. List work as filmography rows: a didone title, one or two plain sentences, an inline amber link, thin rules between rows. Include one contained photographic plate with a monospace caption, one italic didone pullquote between crimson rules, and an end-credit colophon of centered lines on the bare ground.
No neon glow, no purple gradients, no glass cards, no pill radii, no fake player chrome, no readable text inside images. Hover states on links near 150ms with a prefers-reduced-motion guard.
Beaver Builder AI takes it whole; so do v0, Lovable, Figma Make, and Claude Code.
StingerField NoteTC 00:04:10
This site already runs as a one-person crew, with models filling the other chairs. FLUX 3 makes the same bet about the film set, robot arm included.