Overview
What Fish Pond is and what each route shows.
Fish Pond is a set of browser experiments about how an artificial fish can decide what to do. It is built with Next.js, TypeScript, Three.js and Triforge shader graphs, and runs entirely in the browser: there is no backend, no account and no tracking.
| Route | What it is | Interaction |
|---|---|---|
| / | Landing page over the live underwater scene | Open an experiment; click the water to feed the fish |
| /arena | Cognitive Arena: one fish, instinct versus deliberation (scripted, no AI) | Place stimuli, pause, change speed, inspect the fish mind |
| /benchmark | Visual Telemetry Benchmark: three scripted controllers, identical inputs (no AI) | Steer a threat with the cursor, drop food, run the test cycle |
| /underwater | A realistic goldfish scene whose fish are driven by Laya-AI, with a full settings sidebar | Enable Laya-AI, feed, fly with WASD and arrows, tune about 60 graphics settings |
The experiments share one visual language (Liquid Glass overlays over a WebGL canvas), one set of design tokens and one set of coding standards. Everything described here is implemented in the repository; where something is a simulation rather than a real model, this page says so.
Cognitive Arena
Instinct (System 1) versus deliberation (System 1+2) in one fish. Scripted, no AI model.
The arena places a single fish in a wide pond (32 by 16 units, 7 tall). You place stimuli; the fish either reacts instantly from instinct or investigates slowly. The point is to make the difference between the two modes of thought visible.
Stimuli
Eight stimuli are available, each with a hotkey. Two are instant (they fire as soon as you choose them); the rest are placed by clicking the water.
| Key | Stimulus | Class |
|---|---|---|
| 1 | Looming shadow | Known threat (System 1) |
| 2 | Glass tap (instant) | Known threat (System 1) |
| 3 | Light flash (instant) | Known threat (System 1) |
| 4 | Food pellet | Known food (System 1) |
| 5 | Novel rock | Unknown (System 1+2) |
| 6 | Novel food | Unknown (System 1+2) |
| 7 | Drifting leaf | Unknown (System 1+2) |
| 8 | Lure | Unknown (System 1+2) |
System 1: instinct
System 1 only reacts to known stimuli. A reaction latency of 80 to 120 ms passes, then a committed reflex fires. The reflex is chosen by a weighted lottery over designer presets (the same idea as benchmark Case 2). Threat reflexes always pre-empt food strikes, and hunger or fear gate whether a food strike happens at all.
System 1+2: deliberation
Unknown objects trigger a slow, per-object appraisal with five phases: notice, cautious approach, inspect, probe, verdict. The duration scales with novelty, from about 0.7 s for something the fish remembers up to about 4.5 s for something completely new. The verdict (food, safe or threat) is written to short-term memory with a confidence value. Deliberation is a deterministic simulation, so the same inputs always give the same result; it is not a language model.
The arbiter
- System 1 always wins. A reflex aborts any running appraisal.
- After an override, System 2 is locked out for 2.5 seconds, and stays blocked while fear is above 0.45.
- An override adds a valence penalty (0.35) to the interrupted object, so a harmless rock interrupted by a shadow can be judged a threat later. This is intentional: it models a learned association.
Fish state and memory
Hunger starts at 0.45 and rises 0.012 per second; eating relieves 0.5. Fear peaks at 1.0 and decays 0.3 per second. Short-term memory keeps each appraised kind for a span (default 30 s, adjustable from 5 to 120 s with the slider); meeting it again refreshes the timer, and when it expires the kind is novel again.
Scenarios and controls
Four scripted scenarios reproduce the interesting cases: Interrupted inspection, Hunger versus fear, Novelty meets threat, and Learned aversion.
- Keys 1 to 8 choose a stimulus, click the water to place it, Esc cancels.
- Space pauses, R resets, and the speed menu offers 0.5x, 1x and 2x.
- The Stimuli button opens the palette and scenarios pane. On phones the same content lives in tabbed bottom sheets.
Visual Telemetry Benchmark
Three scripted controllers, one set of inputs, compared live. No AI model.
The benchmark runs three fish in three side-by-side ponds. A shared telemetry bus feeds them identical inputs, so any difference in behaviour comes only from the controller.
| Case | Idea | How it behaves |
|---|---|---|
| 1 · System 0, programmed reflex | Binary threshold tree | States IDLE, DART and APPROACH. A threat inside 7 units triggers a dart; food inside 10 units triggers an approach. A dart holds for at least 0.9 s. Transitions are instant, so motion snaps. |
| 2 · Preset lottery (scripted) | Weighted lottery over designer presets | Presets Glide, Startle-Dart, Curious-Hover and Anxious-Freeze are chosen by probability, then blended with smoothing so the fish keeps momentum. |
| 3 · Blended intent (scripted) | Continuous multi-axis blending | No presets. Threat and food intents (ranges 14 and 16 units) drive spine curve, tail frequency and fin resistance continuously. When threat and food are both strong, an ambivalent hesitation appears. |
Inputs
- Cursor: move over any pond to steer a descending hand threat; click a pond or press Drop food to add a pellet.
- Test cycle: a 16-second scripted sequence (approach, then a hand-and-food conflict, then the remainder) piped to all three ponds at once.
Telemetry
Each pond overlays its state, frames per second, tail frequency, spine curvature, tension and a probability bar. All three controllers are clamped by the same physical safety limits (tail 0.8 to 7.5 Hz, spine curve 0.05 to 0.95, fin resistance 0.1 to 1.0, speed 0.2 to 8). Randomness is seeded, and every fish eats a shared pellet independently so the comparison stays fair.
The Compare button opens a matrix that summarises the decision mechanism, transition fidelity, conflict behaviour and performance of each case.
Underwater scene
How the realistic scene is shaded, lit and graded.
The scene is drawn with Three.js, but every custom material is a Triforge node graph: there is no hand-written GLSL in the project. Graphs cover the sandy seabed, rocks, seagrass, the backdrop dome, the mirror-like water surface seen from below, the light shafts and the contact shadow under each fish.
The water column
Light is absorbed per colour channel with a Beer-Lambert model (red dies first, blue survives) and scattered toward the water colour with distance, so far surfaces dissolve into the fog. Absorption and water colour are adjustable in the settings sidebar.
Caustics, shafts and surface
- Caustics: two animated ridged-noise layers multiplied together, faded with depth and with surfaces facing away from the sun.
- Light shafts: additive billboards along the sun direction with streaked noise, up to 24 of them.
- Surface: seen from below, a bright Snell window straight up and total internal reflection beyond it.
- Dappling: the sun light flickers through three sine waves so moving light reaches the fish.
- Seagrass and snow: blades are bent on the CPU each frame in one draw call; up to 1,800 marine-snow particles drift downward.
Post-processing
A Triforge compositor chain applies bloom, colour balance (lift and gain per channel), hue and saturation, a vignette and film grain. Its render targets are multisampled to avoid edge aliasing.
Adaptive quality
The engine samples frame time. If frames average over 27 ms it lowers the pixel ratio by 0.25 (between 1 and 1.5), so weaker GPUs stay fluid.
Fish behaviour and Laya-AI (underwater)
How an AI model, not code, decides what each goldfish does.
The scene holds one large Jikin goldfish and two smaller Tosakin goldfish, loaded once from real GLB models and cloned. Their decisions come from Laya-AI, an open System 1 decision model that runs in your browser through ONNX Runtime Web. Switch it on with the Enable Laya-AI chip. With it off, the fish only drift: there is no scripted fallback that decides for them.
What the model decides
| Decision | What the model is asked | What happens |
|---|---|---|
| Flee | Yes or no: "A human hand is next to the goldfish", given a plain description of where the hand is | At 50% or more the fish flees away from the cursor |
| Eat | Yes or no: "The fish should go and eat the food now", given the fish's first-person view | At 50% or more the fish swims to the nearest pellet |
| Where to swim | Choice between four candidate spots, each described in words (distance, cover, company, depth) | The chosen spot becomes the next destination |
What stays in code
Code only carries out what the model decided and keeps the world physical: steering and speed, the direction that points away from the cursor, personal space between fish, the seabed and rocks, the tank walls, body pose, and swallowing a pellet that reaches the mouth. It never chooses to flee, eat or travel. A question is only asked when it applies: the flee question while a hand is near and the eat question while food is near, and a fish with a hand near is asked at once, nearest first, ahead of everything else.
Temperaments
| Fish | Temperament | Effect on movement |
|---|---|---|
| Jikin | Calm elder | Slow, wanders gently |
| Tosakin | Bold forager | Fast, wanders more |
| Tosakin | Shy follower | Slightly slower, wanders steadily |
The temperament label is also given to the model when it picks a destination.
How fast it is
A flee decision takes about 0.7 s on a desktop with hardware graphics, and the first fish reacts 1.1 to 1.9 s after the pointer enters the scene (measured on a 16-thread laptop); eat and destination decisions take about 1.2 s. The model runs in a Web Worker, so the scene stays smooth, and a fish keeps acting on its last answer until the next one arrives. The first visit downloads about 524 MB of model weights; the browser keeps them, so later visits start in seconds. Slow or small devices can take much longer or fail to start, in which case the fish stay idle and the chip says so.
The wording is part of the experiment
Laya is a text classifier trained on business workflows such as support tickets, invoices and security incidents. It is not trained on fish. Testing it on this task showed what it can and cannot do:
- It answers "should the fish eat?" reliably (86 to 93% when food is in view, near 0% when it is not).
- It cannot grade a danger by distance or speed. Asked "is the fish in danger?" it gave about the same answer for a hand 10 cm away and one 3 m away, with both published checkpoints. The only wording that separated them put the word "danger" into the situation text, which would have been code making the decision, so it was not used.
- It does recognise a plain description of closeness. "A human hand is in the water right next to the goldfish" scored 76 to 80%, and "far across the tank" scored 3 to 31%. So the situation is described in place words, and the model decides on those.
- Small changes in wording change the answer a lot, so the exact sentences are fixed in
fish-perception-text.ts.
These are limits of the model, not tuning that code covers up. When the model disagrees with what you would expect, that is the experiment working.
Food and collisions
A click drops a cluster of 26 pellets (instanced, drag-limited sinking, slight drift). Settled pellets dissolve after 34 seconds, and pellets on the ground trigger a head-down nibble with a bobbing gulp. Fish collide with the seabed and rocks using nose, middle and tail samples, and push apart from each other.
The live readout
The bottom-right text on the underwater page shows, for each fish, a unique marker (#1, #2, #3), its state, appetite, panic and the latest Laya answers (flee and eat, as percentages), updated four times a second. A dash means the model has not answered yet. Hover a fish, or its block, and a small arrow label with the marker, name and state follows the fish for about four seconds while its block scales up. The chip bottom-left shows the model status and the time per decision.
Controls
Mouse, keyboard and touch on every route.
| Input | Action |
|---|---|
| Mouse move | Camera parallax; the cursor is the hand Laya-AI judges (when enabled) |
| Hover a fish or its readout block | An arrow label with the fish's #marker, name and state follows it for about 4 seconds and its readout block scales up |
| Enable Laya-AI chip | Downloads and starts the model (about 524 MB, kept by the browser); click again to switch it off |
| Click or tap the water | Drop fish food |
| W A S D | Move forward, left, back, right |
| Arrow keys | Turn left and right, look up and down |
| Touch pads | Move (bottom-left) and Look (bottom-right) on phones and touch screens |
| Hide HUD | Hides every overlay for a clean view |
Navigation is added on top of the automatic camera, so mouse behaviour is unchanged. The camera stays inside the tank and above the seabed, and the keys are ignored while a slider, number box or menu has focus so the sidebar stays usable from the keyboard.
The arena and benchmark use their own controls (see those sections), and all three experiments have a Hide HUD button.
Scene settings and the JSON workflow
Tune about 60 graphics values, export them, make them the defaults.
The Scene settings sidebar on /underwater exposes 62 graphics values in ten tabs: Water, Light, Caustics, Shafts, Snow, Camera, Grade, Plants, Shadow and Surfaces. Each has a slider and a typed value box; colours use a picker. Changes apply in one of three ways.
| Mode | What happens | Examples |
|---|---|---|
| Live | Applied at once | Exposure, fog density, sun and ambient light, camera, snow, seagrass sway |
| Materials | Triforge graphs are recompiled after a 250 ms pause | Water colour, absorption, caustics, shaft intensity, bump and caustic gain |
| Grade | The compositor is rebuilt after a 250 ms pause | Bloom, vignette, grain, lift, gain, saturation |
Making values the new defaults
- Tune until you are happy, then press Export JSON. A file named
fish-pond-underwater-settings.jsonis downloaded. - Paste its
settingsobject overSCENE_DEFAULTSinsrc/simulation/underwater/underwater-settings.ts. - Rebuild. Your browser’s saved tuning is ignored automatically once the defaults change, so nothing stale overrides them.
Tuning is remembered in the browser (local storage) between visits. Reset all returns to the shipped defaults.
Design system
Liquid Glass tokens, rules and the CSS architecture.
The interface follows a house look called Liquid Glass: frosted, translucent panes floating over the canvas, white text in two weights, hairline dividers and pill controls. Nothing is saturated except the scene behind it.
Rules
- Glass only appears over imagery or the canvas; on flat colour a plain surface token is used.
- At most two stacked glass layers, and the inner layer is a plain translucent fill with no extra blur.
- Text contrast is checked against the worst background pixel, with a stronger tint as the fallback.
- Fallbacks ship with every glass surface: a stronger fill without blur when backdrop-filter is unsupported or transparency is reduced, and no transform transitions when reduced motion is requested.
- Every control has a visible focus ring and a hit area of at least 40 px.
- State is driven by
data-*attribute values, never by toggling classes.
Tokens and naming
All design tokens live in one file, src/app/variables.css, prefixed --fp-. Custom classes carry the signature fp-, with a-ts suffix when TypeScript touches them. Layout and spacing come from Strata CSS utilities first; custom CSS covers surfaces, typography, colour, pseudo-classes and state. Properties are ordered alphabetically, there is no !important, and media queries are range-based.
Transitions
Pages cross-fade over 400 ms with React view transitions, scenes fade in when ready, and the HUD eases in and out with the same duration in both directions.
Accessibility
Keyboard, screen readers, motion and transparency preferences.
- Every page has one h1 and a sequential heading order; overlays use landmarks such as header, nav, aside and section.
- Every icon-only button has an accessible name; decorative icons are hidden from assistive technology.
- Canvases have text alternatives; the landing backdrop is hidden from assistive technology because it is decorative.
- The custom dropdown supports arrow keys, Enter, Escape and outside click, and exposes listbox roles.
- The arena is fully operable from the keyboard; the benchmark cursor threat has no keyboard equivalent yet (see Known limits).
- Reduced motion removes transform transitions and animations; reduced transparency swaps blur for a stronger solid tint.
Performance
What costs what, and how it is kept in check.
- One WebGL engine per route, destroyed on navigation; scenes pause while scrolled out of view.
- The landing page reveals the water as soon as it is drawn and loads the two goldfish models (about 12 MB) afterwards in the background.
- React state from telemetry is throttled (10 Hz in the benchmark, 4 Hz for the fish readout); the simulations themselves run every frame.
- Layout is stable while data updates (fixed-height rows, tabular numerals), so nothing resizes the canvases during play.
- Adaptive pixel ratio steps resolution down on slow frames. Backdrop blur costs GPU time, so glass panes stay small.
Everything was checked in headless Chrome with software rendering; real-GPU frame rates still need measuring on your devices.
SEO and sharing
Metadata, structured data, sitemap and social previews.
- Unique title (50 to 60 characters) and description (120 to 155 characters) per route, canonical URLs, Open Graph and Twitter cards with a 1200 by 630 preview image per route.
- JSON-LD for the website, a WebApplication per experiment, breadcrumbs and an experiment list; a TechArticle on this page.
- A generated sitemap and robots file, a web app manifest, favicon set and apple touch icon, plus an llms.txt summary for AI crawlers.
- The site address comes from
NEXT_PUBLIC_SITE_URL, so every absolute URL follows your domain.
Tech stack and structure
Libraries and where things live.
| Area | Technology |
|---|---|
| Framework | Next.js 16 (App Router), React 19, TypeScript |
| 3D | Three.js, Triforge shader-core and compositor-core |
| Styling | Strata CSS utilities plus CSS Modules and design tokens |
| Fish AI | Laya-AI (ONNX, int8) through ONNX Runtime Web in a Web Worker, tokenised with Hugging Face Tokenizers |
| Hosting | Vercel (static pages, no server code); model weights are fetched from Hugging Face by the browser |
src/
app/ routes, global CSS, tokens, sitemap, robots, manifest
components/ shared UI: HUD icons, glass select, landing backdrop
arena/ Cognitive Arena panels and scene
underwater/ stage, canvas, settings panel, touch pads, readout
config/ site constants and SEO builders
simulation/
arena/ System 1, System 2, arbiter, memory, scenarios
controllers/ benchmark Cases 1 to 3
underwater/ engine, Triforge graphs, fish motor control, settings
ai/ Laya runtime, worker, perception wording, fish intent controller
types/ shared types for arena, benchmark and underwater
scripts/ copy-ort.ts (copies the ONNX Runtime wasm into public/ort)
public/ goldfish models, icons, social preview imagesRun, build and deploy
Commands and hosting.
git clone --recurse-submodules https://github.com/AftabIbrahimKazi/fish-pond.git
cd fish-pond
npm install
npm run dev # http://localhost:3000
npm run build # production build
npm run start # serve the buildThe site is a static Next.js build and deploys to Vercel with no configuration. Set NEXT_PUBLIC_SITE_URL to the final address so canonical URLs, the sitemap and social previews point at it. npm install copies the ONNX Runtime wasm files into public/ort, and the site sends cross-origin isolation headers (needed for threaded WebAssembly). The model weights are not part of the deployment: the browser fetches them from Hugging Face, so the hosting plan only serves the small app.
Known limits and roadmap
What is not finished, stated plainly.
- Not yet tested on a real GPU: frame rates, the automated benchmark cycle and the arena need measuring on real hardware.
- The benchmark cursor threat has no keyboard equivalent.
- The environment reflection map is baked once at load, so changing the water colour does not update reflections until a reload.
- Rocks have no collision for the free camera.
- Fish behaviour, terrain shape and scenery placement are not yet adjustable in the settings sidebar.
- Laya-AI is a general text classifier, not a fish model. It cannot grade danger by distance, so flee decisions rest on plain place words (see the wording notes above).
- A decision takes about a second per fish and the model needs a large download and a lot of memory while loading (estimated at 1 to 2 GB, not yet measured). Phones and low-end laptops may not manage it.
- Without Laya-AI enabled the underwater fish only drift. The landing page backdrop runs the same scene without the model, so its fish drift too.
- Fish do not school: grouping would have to come from the model choosing destinations near each other.
- The Arena and Benchmark are scripted. A second model (ReasonLite, a reasoning model) for System 2 is not implemented.
- Next on the list: System 2 reasoning for the underwater fish, more feeding animation, and touch support for feeding in the benchmark.
Credits
Built on open tools.
Fish Pond is written by Aftab Ibrahim Kazi. It stands on Next.js, React, Three.js, the Triforge shader and compositor packages and Strata CSS. The goldfish models ship in the repository; confirm their licences before redistributing them elsewhere. The source is on GitHub.
Laya-AI
- Laya (the decision model) was created by Nandakishor M and Convai Innovations: convaiinnovations/laya (Apache-2.0).
- The 8-bit browser build and the runtime that Fish Pond's Laya code is ported from were made by nvkudva: github.com/nvkudva/laya-web and nvkudva/laya-web-q8 (Apache-2.0 weights). The laya-web repository does not state a licence for its code; it is used here with credit, and will be removed or replaced if its author asks.
- Also used: ONNX Runtime Web by Microsoft and Tokenizers by Hugging Face.