Engineering note
What actually broke when we added real-time collaboration
Collaborative edits and live Firebase updates could change the same document, and tasks started appearing twice.
Adding real-time collaboration to the product I was working on meant dealing with more than people editing together. The application could change the document too.
The editor used Tiptap, Yjs, and Hocuspocus for collaborative editing, while Firebase Realtime Database supplied live application data. Updates from that data could change task content or reorder parts of the document.
Both needed to work in the same editor. When their updates overlapped, tasks and user-content nodes could appear twice.
Two ways to change the same document
One update path came from people editing collaboratively. The other came from the application responding to Firebase data.
The intended behaviour sounded straightforward: people should be able to keep editing while live data updated the relevant content and ordering. A task should still appear once, regardless of which update arrived first.
But those paths weren’t independent. They both modified the same document, so getting each integration working on its own wasn’t enough. Their interaction needed coordination too.
- People editingCollaborative editTiptap, Yjs, Hocuspocus
- The applicationFirebase updateTask content and ordering
- One documentSame task, twiceWhen both paths overlap
Tasks started appearing twice
The issue showed up during multi-user editing and Firebase-driven reordering. Tasks and user-content nodes could appear more than once.
The cause was a race condition between the update paths. Each could complete its own work, but overlapping operations could leave duplicate nodes in the document.
A simplified way to picture it is a Firebase event triggering a document update while a collaborative update arrives. Both paths modify the document, and the resulting content contains duplicates. That’s the failure pattern, rather than an exact trace of the callbacks involved.
Using Yjs didn’t automatically prevent this. The application still decided what document changes to generate from external data. Two insertions could represent the same logical task, even if the collaborative system could synchronize both operations.
The problem was in how those update paths interacted. I hadn’t established a defect in Yjs itself.
The fixes that didn’t quite hold
I tried update flags, deduplication sets, and debouncing. I also worked through provider synchronization, lifecycle checks, and cleanup around update handling.
| Attempt | What it changed |
|---|---|
| Update flags | Reduced the symptoms, but didn’t prevent the overlap. |
| Deduplication sets | Tracked content already handled, but didn’t solve the coordination producing the duplicates. |
| Debouncing | Made close-together updates rarer, but couldn’t guarantee they wouldn’t overlap. |
| Provider synchronization, lifecycle checks and cleanup | Reduced the symptoms, but didn’t prevent the overlap. |
Those changes reduced the symptoms, but they didn’t fully prevent the problematic overlap.
Debouncing, for example, could reduce how often updates ran close together. It couldn’t guarantee that competing operations wouldn’t overlap. Deduplication could track content already handled, but it didn’t fully solve the coordination problem producing the duplicates.
That distinction mattered. I needed to control how the affected operations executed, rather than keep reducing the chances of them colliding.
Coordinating the updates with a mutex
The fix was to protect the relevant editor-update paths with a mutex.
An operation acquired the lock, completed its protected update, and released it. Another operation using that same lock had to wait before entering the protected section.
One operation at a time
- Acquire the lock
- Complete the protected update
- Release it
That gave the competing operations a controlled execution order within those paths. They could still update the document, but the affected work could no longer overlap.
The useful part wasn’t the mutex by itself. It was putting the coordination around the work that needed it. A lock only helps when the competing operations share it and respect its boundary.
What the fix proved, and what it didn’t
After coordinating the affected update paths, the duplication was resolved in the multi-user scenarios tested.
The tradeoff was waiting. Competing updates could now have to wait for the protected operation to finish before proceeding. That was acceptable for preventing the duplicate-content behaviour, though I don’t have timing measurements to quantify the cost.
The result also had a boundary. It covered the protected operations and the scenarios tested. A local mutex doesn’t automatically coordinate every browser, server, or database writer.
The observed issue was duplicated editor content. There wasn’t evidence here to claim lost user text, database corruption, or a fix that guaranteed consistency across the entire system.
Collaboration includes the updates your app generates
The difficult part was making collaborative edits coexist with changes generated by the application. Both could modify the same document, and their interaction deserved as much attention as either integration on its own.
The earlier attempts helped with the symptoms. The mutex addressed the overlap between the affected operations.
My takeaway was concrete: when content appears twice in a collaborative editor, trace every path that can modify the document. Include the background updates, external-data listeners, and reordering logic alongside direct user edits. Then identify which operations need to be coordinated, and where that coordination actually applies.