Problems I didn't know how to solve when I started.
02 · strata ·
Where should the information for skipping files live?
the problem
A search should open only the log files that can hold the answer. Knowing which ones means storing each file's time range and word summary somewhere, and reading it costs time on every search.
what I considered
Keep the summary inside each file and read just that part from storage. A separate database like Postgres. A SQLite file inside the server.
what I picked
A SQLite file with one row per log file: time range, size, and its bloom filter. A search reads it first and opens only what it can't rule out.
what I'd change
On 196 files, a one-hour search for a common word opened 2 or 3 files in 55 ms. Across all time it opened 171 in 2.1 s. The cost: one file the server can't lose.
What should happen when two devices edit the same file?
the problem
Two devices can edit the same file while one is offline. If the last upload just wins, one of those edits disappears and nobody is told.
what I considered
Last upload wins, locking files, merging edits automatically, or keeping both versions.
what I picked
Keep both. The server turns away any upload based on an old version, and the device saves its own edit as a conflicted copy next to the original.
what I'd change
My first prototype let the server pick the version, so it never caught this. There's still a tiny window where a save during a download can get overwritten.