diff --git a/README.md b/README.md
index abb1feb..46bc9de 100644
--- a/README.md
+++ b/README.md
@@ -30,6 +30,7 @@ go install .
```
midden init Create the vault, write a vault README and gitignore.
midden ingest ics calendar.ics Backfill past events from a calendar export.
+midden ingest git ~/src/project Backfill what you were working on, from commit history.
midden add "text" Append an entry to today.
midden add "text" --tag work Append with tags.
midden today Open today's day file in the editor.
@@ -59,8 +60,29 @@ midden import - --date 2026-06-10 Read stdin and file it on a chosen
midden audio Record a voice memo and append it to today.
midden audio --duration 30s --transcribe Record for 30s then transcribe with OpenAI Whisper.
midden ingest ics calendar.ics Append calendar events from an .ics export.
+midden ingest ics calendar.ics --since 2015-01-01 Narrow the range to ingest.
+midden ingest git ~/src/one ~/src/two Append commit history from local repositories.
+midden ingest git ~/src/work --author me@example.com --stat
```
+Ingest reads the whole export by default, because backfilling years of calendar history is the point
+of it. Recurring series are expanded into the occurrences they actually produced, so a weekly one-to-one
+running since 2019 contributes every week rather than a single event in 2019. `EXDATE` cancellations are
+honored and instances the calendar moved replace the occurrence they override, so a rescheduled meeting
+appears once at its real time rather than twice. Occurrences already in the vault are skipped, so running
+the same import twice changes nothing.
+
+Rules using `BYSETPOS`, `BYYEARDAY`, or `BYWEEKNO` are not expanded on those parts, and a frequency
+outside daily, weekly, monthly, and yearly is not expanded at all. Ingest counts and reports both cases
+rather than passing off a partial calendar as a complete one.
+
+`ingest git` appends one entry per commit across any number of local repositories. A calendar says where
+you were; commit history says what you were working on, and the two together reconstruct a working life
+far better than either alone. Commits are filed by author date, so rebased or cherry-picked work still
+lands on the day it was written. Merge commits are skipped unless `--merges` is given, `--author` narrows
+a shared repository to your own commits, and `--stat` adds changed-file and line counts at the cost of a
+diff per commit. Commits already in the vault are skipped by hash.
+
@@ -95,8 +117,123 @@ midden tag work List entries with a tag.
midden tags Show the tag histogram.
midden recall "token rotation strategy" Semantic search over indexed entries.
midden chat "when did I last see Mom?" Ask an LLM a question using recalled entries as evidence.
+midden chat --since 30-days-ago "what did I do?" Answer from every entry in a date range.
+midden chat --sweep "what do you know about my life?" Answer from the whole vault.
+midden weave --tag calendar Show what recurs, when it started, and when it stopped.
+midden ask -i Answer a question about a gap in your own record.
midden reindex Build the embedding index used by recall.
+midden reindex --full Re-embed everything, needed only after changing provider.
+```
+
+Reindex reuses the vector it already has for any entry whose text has not changed, so rebuilding after
+adding a day costs one provider call rather than re-embedding the whole vault. Long rebuilds checkpoint
+as they go and each provider call has its own deadline, so an interrupted backfill resumes from where it
+stopped instead of throwing away the embeddings it already paid for.
+
+`recall` and `chat` both accept `--since` and `--until`, which take any date `midden` understands
+(`2024-03-01`, `30-days-ago`, `monday`). Scoping matters because ranking by similarity alone answers
+"what did I write about X" well and "what happened last March" badly: the closest matches to a question
+about a period are often entries from other periods. Giving `chat` a range makes it read every entry in
+that range instead of the closest few, summarizing in chunks when the range is too large to read at
+once. `--sweep` does the same across whatever is in scope, which is the whole vault by default.
+
+Every `chat` answer also carries a summary counted over every indexed entry in scope: how many entries,
+what span they cover, the tag histogram, and entries per month. Questions about the shape of the record
+are answered from those counts rather than from a handful of retrieved entries.
+
+### Scheduled entries
+
+A vault holding an imported calendar contains appointments that have not happened yet, which quietly
+breaks anything meaning "latest". `last` and `recent` therefore stop at now, and `stats` reports what is
+booked ahead separately from the span of what actually happened:
+
```
+Span
+ First: 2022-08-18 07:00:00
+ Last: 2026-08-26 15:10:23
+ Ahead: 264 scheduled, through 2027-03-12
+```
+
+Pass `--future` to `last` or `recent` when you do want what is coming. `undo` never touches a scheduled
+entry: it removes the last thing you wrote, and nothing you wrote lives in the future.
+
+`streak` counts only days you actually wrote something. A backfilled vault has entries on thousands of
+days the person never wrote a word, and a streak counted over imported events would congratulate you for
+appointments you merely attended.
+
+### Ask
+
+Backfill has a ceiling, and it is worth being plain about where it sits. Calendars record where you were
+scheduled. Commit logs record what you shipped. Both are projections of a life rather than the life, and
+neither carries what you thought or decided, because nothing recorded that at the time. No further import
+fixes this: for anyone who was not already keeping a journal, that material does not exist to import.
+
+`midden ask` closes the gap the only way it can be closed. It reads what weave computed, finds a place
+the record proves something is missing, and asks about it:
+
+```
+William- Martial arts stopped. What happened?
+ 216 times over 2.6 years, ending 2025-08-07. Nothing since, 1.0 years ago.
+```
+
+Answer it and the reply becomes an ordinary entry, with your words leading and the question trailing as
+attribution, so the record shows what you said rather than what midden asked. That question is never
+asked again. `--skip` dismisses one for good, because a queue that keeps returning a question you have
+already rejected teaches you to stop reading it.
+
+Questions are only raised about things worth explaining. A commitment that ran for months and stopped
+qualifies; a school-year reminder repeated for one term does not, and neither does anything merely
+between seasons. Questions come
+from arithmetic over the record, never from a model, so nothing is asked about something that did not
+happen. Threads still running are never asked about at all: frequency alone cannot tell a commitment that
+mattered from a chore that recurred, and a vapid prompt teaches you to ignore the next one.
+
+An answer is the first thing in a vault that no import could have produced.
+
+### Weave
+
+Search answers what you already know to ask about. `midden weave` answers what you cannot ask.
+
+A person can recall what they did but cannot perceive absence, because nothing marks the last time
+something happened. Weave groups the record into threads, measures each one's own cadence, and reports
+which have gone quiet for far longer than their rhythm allows. It also finds handoffs, where one thread
+ended and another began soon after, and crossings, the days where separate sources both recorded
+something and so can say what neither says alone.
+
+Threads are grouped by the meaningful words in a title, which survives the drift of a handwritten
+calendar without merging things that are genuinely different. A word added to a short title is a subject,
+not noise: "Doctor appt" and "Hannah doctor appt" overlap heavily but the second says whose appointment
+it was, and folding them together would treat two people's appointments as one thread and then report a
+gap spanning the distance between two unrelated lives. `--explain` lists the headlines behind every
+thread, because a claim that something ended rests entirely on what was grouped together, and that has to
+be checkable before it is believed.
+
+A thread that has gone quiet is not automatically over. Something that has already come back from a gap
+this long before is between seasons, not finished, and weave says so rather than announcing an ending
+that never happened. A spring show silent in August has simply not come round yet.
+
+Silences are found per source as well as overall, because one loud source can flood the months where
+another went quiet: commits pouring in during years the calendar recorded nothing would otherwise hide
+exactly the silence worth asking about. A silence found in the whole record is never repeated per source,
+and ask raises at most one silence per source, the longest.
+
+It also finds silences: stretches where the record itself went quiet, bounded by activity on both sides
+so the start and end of a record are never mistaken for holes. A thread ending is one commitment
+stopping. A silence is the record failing, which is both larger and completely invisible from inside it,
+since a person notices a class ending but never that years went unrecorded.
+
+Every figure is counted rather than inferred. There is no model in the detection path and nothing to
+invent. Threads are grouped by the meaningful words in a title rather than the title itself, because a
+handwritten calendar records one standing arrangement under many spellings, and grouping on the exact
+string splits it into fragments that each appear to end whenever the wording drifts.
+
+Use `--tag calendar` to weave a life source on its own. Commit history repeats boilerplate subjects
+across repositories, which crowds out real threads.
+
+`--context-chars` sets how much entry text goes to the model in one call. The default suits a model with
+a large context window. A small local model needs a much lower value, because it spends minutes on a
+prompt a hosted model reads in seconds, which makes a sweep look like a hang. Try `--context-chars 6000`
+against a 3B local model and raise it from there.
diff --git a/cmd/chat_context.go b/cmd/chat_context.go
new file mode 100644
index 0000000..44cdb9a
--- /dev/null
+++ b/cmd/chat_context.go
@@ -0,0 +1,209 @@
+package cmd
+
+import (
+ "context"
+ "errors"
+ "fmt"
+ "strings"
+ "sync"
+
+ "github.com/spf13/cobra"
+
+ "github.com/dcadolph/midden/index"
+ "github.com/dcadolph/midden/llm"
+)
+
+// defaultContextChars is how much entry text is sent to the model in one call,
+// measured in characters. A range that fits is read whole; a larger one is
+// summarized in chunks of this size so the answer still covers the whole span
+// rather than a prefix of it.
+//
+// The default suits a model with a large context window. A small local model
+// needs a far lower value: at this size it spends minutes per chunk, which turns
+// a sweep into an apparent hang. That is what --context-chars is for.
+const defaultContextChars = 400000
+
+// chatChunkWorkers bounds how many chunk summaries run at once.
+const chatChunkWorkers = 4
+
+// chunkSystem frames the map step of a swept range. The summaries are read only
+// by the reduce step, so they are told to keep the specifics an answer needs
+// rather than to read well on their own.
+const chunkSystem = "You are compressing part of a personal journal and calendar archive so a later question " +
+ "can be answered from it. Preserve concrete specifics: dates, people, places, projects, events, " +
+ "recurring commitments, and anything that marks a change. Drop routine filler. " +
+ "Do not speculate beyond the entries and do not add a preamble."
+
+// sweepContext renders every entry in the range as chat context. Ranges within
+// the budget are rendered verbatim; larger ranges are summarized in
+// chronological chunks first, which keeps coverage of the whole span at the
+// cost of detail rather than truncating the span at full detail.
+func sweepContext(
+ ctx context.Context,
+ cmd *cobra.Command,
+ chat llm.Chatter,
+ question string,
+ entries []index.Entry,
+ contextChars int,
+) (string, error) {
+ if len(entries) == 0 {
+ return "", nil
+ }
+ if entriesSize(entries) <= contextChars {
+ return renderEntries(entries), nil
+ }
+ chunks := chunkEntries(entries, contextChars)
+ fmt.Fprintf(cmd.ErrOrStderr(),
+ "Range holds %d entries, too many to read at once; summarizing in %d chunks.\n",
+ len(entries), len(chunks))
+ summaries, err := summarizeChunks(ctx, cmd, chat, question, chunks)
+ if err != nil {
+ return "", err
+ }
+ return strings.Join(summaries, "\n\n"), nil
+}
+
+// summarizeChunks compresses each chunk concurrently and returns the summaries
+// in chronological order. Any chunk failing fails the sweep, because a silently
+// dropped chunk would leave a hole in a range the answer claims to cover.
+func summarizeChunks(
+ ctx context.Context,
+ cmd *cobra.Command,
+ chat llm.Chatter,
+ question string,
+ chunks [][]index.Entry,
+) ([]string, error) {
+ out := make([]string, len(chunks))
+ errs := make([]error, len(chunks))
+ sem := make(chan struct{}, chatChunkWorkers)
+ // Summaries finish out of order and a slow model can take minutes per chunk,
+ // so completions are reported as they land rather than leaving the user
+ // watching a still screen. The mutex keeps concurrent lines from interleaving.
+ var (
+ wg sync.WaitGroup
+ progress sync.Mutex
+ done int
+ )
+ for i, chunk := range chunks {
+ select {
+ case sem <- struct{}{}:
+ case <-ctx.Done():
+ return nil, fmt.Errorf("summarize range: %w", ctx.Err())
+ }
+ wg.Add(1)
+ go func(i int, chunk []index.Entry) {
+ defer wg.Done()
+ defer func() { <-sem }()
+ label := chunkLabel(chunk)
+ user := "Later question: " + question + "\n\nEntries:\n" + renderEntries(chunk)
+ reply, err := chat.Reply(ctx, chunkSystem, []llm.Message{{Role: "user", Content: user}})
+ if err != nil {
+ errs[i] = fmt.Errorf("summarize %s: %w", label, err)
+ return
+ }
+ out[i] = "--- " + label + "\n" + strings.TrimSpace(reply)
+ progress.Lock()
+ done++
+ fmt.Fprintf(cmd.ErrOrStderr(), "Summarized %d/%d (%s)\n", done, len(chunks), label)
+ progress.Unlock()
+ }(i, chunk)
+ }
+ wg.Wait()
+ if err := errors.Join(errs...); err != nil {
+ return nil, err
+ }
+ return out, nil
+}
+
+// chunkEntries splits chronologically ordered entries into runs whose bodies
+// stay within the budget. An entry larger than the budget on its own becomes a
+// chunk of one rather than being dropped or split mid-body.
+func chunkEntries(entries []index.Entry, budget int) [][]index.Entry {
+ var (
+ out [][]index.Entry
+ chunk []index.Entry
+ size int
+ )
+ for _, e := range entries {
+ n := len(e.Body)
+ if len(chunk) > 0 && size+n > budget {
+ out = append(out, chunk)
+ chunk, size = nil, 0
+ }
+ chunk = append(chunk, e)
+ size += n
+ }
+ if len(chunk) > 0 {
+ out = append(out, chunk)
+ }
+ return out
+}
+
+// chunkLabel renders the date span a chunk covers.
+func chunkLabel(chunk []index.Entry) string {
+ if len(chunk) == 0 {
+ return "empty range"
+ }
+ first := chunk[0].Time.Format(layoutDate)
+ last := chunk[len(chunk)-1].Time.Format(layoutDate)
+ if first == last {
+ return first
+ }
+ return first + " to " + last
+}
+
+// entriesSize returns the total body length across the entries.
+func entriesSize(entries []index.Entry) int {
+ total := 0
+ for _, e := range entries {
+ total += len(e.Body)
+ }
+ return total
+}
+
+// renderEntries renders entries as the dated blocks the chat model reads.
+func renderEntries(entries []index.Entry) string {
+ var b strings.Builder
+ for _, e := range entries {
+ b.WriteString("--- ")
+ b.WriteString(e.Time.Format(layoutDateTime))
+ if len(e.Tags) > 0 {
+ b.WriteString(" [")
+ b.WriteString(strings.Join(e.Tags, ", "))
+ b.WriteString("]")
+ }
+ b.WriteString("\n")
+ b.WriteString(e.Body)
+ b.WriteString("\n\n")
+ }
+ return b.String()
+}
+
+// renderDigest renders the corpus summary supplied with every answer. The
+// counts come from every indexed entry in scope, which is what lets the model
+// answer about the record as a whole rather than about the entries it was sent.
+func renderDigest(d index.Digest, label string) string {
+ var b strings.Builder
+ fmt.Fprintf(&b, "Corpus summary, counted over every indexed entry in scope (%s)\n", label)
+ if d.Entries == 0 {
+ b.WriteString("Entries: 0\n")
+ return b.String()
+ }
+ fmt.Fprintf(&b, "Entries: %d\n", d.Entries)
+ fmt.Fprintf(&b, "Span: %s to %s\n", d.First.Format(layoutDate), d.Last.Format(layoutDate))
+ if len(d.TopTags) > 0 {
+ parts := make([]string, len(d.TopTags))
+ for i, t := range d.TopTags {
+ parts[i] = fmt.Sprintf("%s (%d)", t.Tag, t.Count)
+ }
+ fmt.Fprintf(&b, "Tags by entry count: %s\n", strings.Join(parts, ", "))
+ }
+ if len(d.Months) > 0 {
+ parts := make([]string, len(d.Months))
+ for i, m := range d.Months {
+ parts[i] = fmt.Sprintf("%s (%d)", m.Month, m.Count)
+ }
+ fmt.Fprintf(&b, "Entries per month: %s\n", strings.Join(parts, ", "))
+ }
+ return b.String()
+}
diff --git a/cmd/chat_context_test.go b/cmd/chat_context_test.go
new file mode 100644
index 0000000..45159e3
--- /dev/null
+++ b/cmd/chat_context_test.go
@@ -0,0 +1,148 @@
+package cmd
+
+import (
+ "fmt"
+ "strings"
+ "testing"
+ "time"
+
+ "github.com/google/go-cmp/cmp"
+ "github.com/google/go-cmp/cmp/cmpopts"
+
+ "github.com/dcadolph/midden/index"
+ "github.com/dcadolph/midden/internal/util"
+)
+
+// chunkEntry builds an indexed entry dated d days after 2024-01-01 with a body
+// of the given length.
+func chunkEntry(d, size int) index.Entry {
+ return index.Entry{
+ Time: time.Date(2024, time.January, 1, 12, 0, 0, 0, time.Local).AddDate(0, 0, d),
+ Body: strings.Repeat("x", size),
+ }
+}
+
+func TestChunkEntries(t *testing.T) {
+ t.Parallel()
+ tests := []struct {
+ WantSizes []int
+ In []index.Entry
+ Budget int
+ }{{ // Test 0: Entries that fit the budget stay in one chunk.
+ In: []index.Entry{chunkEntry(0, 10), chunkEntry(1, 10)},
+ Budget: 100,
+ WantSizes: []int{2},
+ }, { // Test 1: The chunk breaks before the entry that would exceed the budget.
+ In: []index.Entry{chunkEntry(0, 40), chunkEntry(1, 40), chunkEntry(2, 40)},
+ Budget: 100,
+ WantSizes: []int{2, 1},
+ }, { // Test 2: An entry larger than the budget becomes a chunk of one rather than being dropped.
+ In: []index.Entry{chunkEntry(0, 10), chunkEntry(1, 500), chunkEntry(2, 10)},
+ Budget: 100,
+ WantSizes: []int{1, 1, 1},
+ }, { // Test 3: No entries produce no chunks.
+ In: nil,
+ Budget: 100,
+ WantSizes: nil,
+ }}
+ for testNum, test := range tests {
+ t.Run(fmt.Sprintf("test %d", testNum), func(t *testing.T) {
+ t.Parallel()
+ chunks := chunkEntries(test.In, test.Budget)
+ got := make([]int, len(chunks))
+ total := 0
+ for i, c := range chunks {
+ got[i] = len(c)
+ total += len(c)
+ }
+ if diff := cmp.Diff(test.WantSizes, got, cmpopts.EquateEmpty()); diff != "" {
+ t.Errorf("chunk sizes mismatch (-want +got):\n%s", diff)
+ }
+ if total != len(test.In) {
+ t.Errorf("chunking dropped entries: want %d total, got %d", len(test.In), total)
+ }
+ })
+ }
+}
+
+func TestChunkLabel(t *testing.T) {
+ t.Parallel()
+ tests := []struct {
+ Want string
+ In []index.Entry
+ }{{ // Test 0: A multi-day chunk renders as a span.
+ In: []index.Entry{chunkEntry(0, 1), chunkEntry(10, 1)},
+ Want: "2024-01-01 to 2024-01-11",
+ }, { // Test 1: A single-day chunk renders as one date.
+ In: []index.Entry{chunkEntry(0, 1)},
+ Want: "2024-01-01",
+ }, { // Test 2: An empty chunk is labeled rather than panicking.
+ In: nil,
+ Want: "empty range",
+ }}
+ for testNum, test := range tests {
+ t.Run(fmt.Sprintf("test %d", testNum), func(t *testing.T) {
+ t.Parallel()
+ if diff := cmp.Diff(test.Want, chunkLabel(test.In)); diff != "" {
+ t.Errorf("mismatch (-want +got):\n%s", diff)
+ }
+ })
+ }
+}
+
+func TestRenderDigestCarriesWholeCorpusCounts(t *testing.T) {
+ t.Parallel()
+ d := index.Digest{
+ Entries: 4213,
+ First: time.Date(2015, time.June, 2, 9, 0, 0, 0, time.Local),
+ Last: time.Date(2026, time.August, 24, 18, 0, 0, 0, time.Local),
+ TopTags: []util.TagCount{{Tag: "calendar", Count: 3900}, {Tag: "work", Count: 210}},
+ Months: []index.MonthCount{{Month: "2015-06", Count: 12}, {Month: "2026-08", Count: 40}},
+ }
+ got := renderDigest(d, "whole vault")
+ for _, want := range []string{
+ "every indexed entry in scope (whole vault)",
+ "Entries: 4213",
+ "Span: 2015-06-02 to 2026-08-24",
+ "calendar (3900)",
+ "2026-08 (40)",
+ } {
+ if !strings.Contains(got, want) {
+ t.Errorf("digest missing %q:\n%s", want, got)
+ }
+ }
+}
+
+func TestRenderDigestEmptyRange(t *testing.T) {
+ t.Parallel()
+ got := renderDigest(index.Digest{}, "since 2030-01-01")
+ if !strings.Contains(got, "Entries: 0") {
+ t.Errorf("want a zero count, got:\n%s", got)
+ }
+ if strings.Contains(got, "Span:") {
+ t.Errorf("want no span for an empty range, got:\n%s", got)
+ }
+}
+
+func TestEntriesSize(t *testing.T) {
+ t.Parallel()
+ got := entriesSize([]index.Entry{chunkEntry(0, 10), chunkEntry(1, 25)})
+ if diff := cmp.Diff(35, got); diff != "" {
+ t.Errorf("mismatch (-want +got):\n%s", diff)
+ }
+}
+
+func TestRenderEntriesIncludesDateAndTags(t *testing.T) {
+ t.Parallel()
+ e := index.Entry{
+ Time: time.Date(2024, time.March, 3, 14, 30, 0, 0, time.Local),
+ Tags: []string{"work", "travel"},
+ Body: "flew to Denver",
+ }
+ got := renderEntries([]index.Entry{e})
+ for _, want := range []string{"2024-03-03 14:30:00", "[work, travel]", "flew to Denver"} {
+ if !strings.Contains(got, want) {
+ t.Errorf("rendered entry missing %q:\n%s", want, got)
+ }
+ }
+}
diff --git a/cmd/chat_sweep_test.go b/cmd/chat_sweep_test.go
new file mode 100644
index 0000000..ae7996f
--- /dev/null
+++ b/cmd/chat_sweep_test.go
@@ -0,0 +1,155 @@
+package cmd
+
+import (
+ "context"
+ "errors"
+ "fmt"
+ "strings"
+ "sync/atomic"
+ "testing"
+
+ "github.com/spf13/cobra"
+
+ "github.com/dcadolph/midden/index"
+ "github.com/dcadolph/midden/llm"
+)
+
+// mockChatter is a Chatter whose reply is produced by a configurable function.
+type mockChatter struct {
+ // ReplyFunc produces the reply for one call, receiving the call ordinal.
+ ReplyFunc func(call int, system string, history []llm.Message) (string, error)
+ // calls counts how many times Reply was invoked.
+ calls atomic.Int64
+}
+
+// Name identifies the mock provider.
+func (m *mockChatter) Name() string { return "mock:chatter" }
+
+// Reply delegates to ReplyFunc, recording the call ordinal.
+func (m *mockChatter) Reply(_ context.Context, system string, history []llm.Message) (string, error) {
+ return m.ReplyFunc(int(m.calls.Add(1)-1), system, history)
+}
+
+// testContextChars is a small context budget, so the sizing tests build a few
+// kilobytes of entries rather than a few megabytes.
+const testContextChars = 3000
+
+// sweepCmd returns a command with output discarded, for driving sweepContext.
+func sweepCmd() *cobra.Command {
+ c := &cobra.Command{}
+ c.SetOut(new(strings.Builder))
+ c.SetErr(new(strings.Builder))
+ return c
+}
+
+func TestSweepContextRendersSmallRangeVerbatim(t *testing.T) {
+ t.Parallel()
+ chat := &mockChatter{ReplyFunc: func(int, string, []llm.Message) (string, error) {
+ return "", errors.New("summarizer must not run for a range within budget")
+ }}
+ entries := []index.Entry{chunkEntry(0, 10), chunkEntry(1, 10)}
+ got, err := sweepContext(context.Background(), sweepCmd(), chat, "what happened", entries, testContextChars)
+ if err != nil {
+ t.Fatalf("sweepContext: %v", err)
+ }
+ if chat.calls.Load() != 0 {
+ t.Errorf("want no summarizer calls, got %d", chat.calls.Load())
+ }
+ if !strings.Contains(got, "2024-01-01") || !strings.Contains(got, "2024-01-02") {
+ t.Errorf("want both entries rendered, got:\n%s", got)
+ }
+}
+
+func TestSweepContextSummarizesOversizeRangeInOrder(t *testing.T) {
+ t.Parallel()
+ // Entries at two thirds of the budget force one chunk each and push the
+ // total well past what fits in a single call.
+ size := testContextChars * 2 / 3
+ entries := []index.Entry{}
+ for i := range 5 {
+ entries = append(entries, chunkEntry(i, size))
+ }
+ chat := &mockChatter{ReplyFunc: func(_ int, _ string, history []llm.Message) (string, error) {
+ // Echo the first date in the chunk so ordering is verifiable.
+ line := strings.SplitN(history[0].Content, "--- ", 2)[1]
+ return "summary of " + strings.SplitN(line, " ", 2)[0], nil
+ }}
+ got, err := sweepContext(context.Background(), sweepCmd(), chat, "what happened", entries, testContextChars)
+ if err != nil {
+ t.Fatalf("sweepContext: %v", err)
+ }
+ if chat.calls.Load() < 2 {
+ t.Fatalf("want the range summarized in several chunks, got %d calls", chat.calls.Load())
+ }
+ // Every chunk must appear, and the summaries must stay chronological even
+ // though they were produced concurrently.
+ var last string
+ for _, line := range strings.Split(got, "\n") {
+ if !strings.HasPrefix(line, "--- ") {
+ continue
+ }
+ date := strings.SplitN(strings.TrimPrefix(line, "--- "), " ", 2)[0]
+ if last != "" && date < last {
+ t.Errorf("summaries out of order: %s came after %s", date, last)
+ }
+ last = date
+ }
+ if last == "" {
+ t.Errorf("want labeled chunk summaries, got:\n%s", got)
+ }
+}
+
+func TestSweepContextFailsWhenAChunkFails(t *testing.T) {
+ t.Parallel()
+ size := testContextChars * 2 / 3
+ entries := []index.Entry{}
+ for i := range 5 {
+ entries = append(entries, chunkEntry(i, size))
+ }
+ chat := &mockChatter{ReplyFunc: func(call int, _ string, _ []llm.Message) (string, error) {
+ if call == 1 {
+ return "", errors.New("provider exploded")
+ }
+ return "fine", nil
+ }}
+ // A dropped chunk would leave a silent hole in a range the answer claims to
+ // cover, so a single chunk failure must fail the sweep.
+ _, err := sweepContext(context.Background(), sweepCmd(), chat, "what happened", entries, testContextChars)
+ if err == nil {
+ t.Fatal("want an error when a chunk summary fails")
+ }
+ if !strings.Contains(err.Error(), "provider exploded") {
+ t.Errorf("want the provider failure surfaced, got %v", err)
+ }
+}
+
+func TestSweepContextEmptyRange(t *testing.T) {
+ t.Parallel()
+ chat := &mockChatter{ReplyFunc: func(int, string, []llm.Message) (string, error) {
+ return "", errors.New("must not run")
+ }}
+ got, err := sweepContext(context.Background(), sweepCmd(), chat, "what happened", nil, testContextChars)
+ if err != nil {
+ t.Fatalf("sweepContext: %v", err)
+ }
+ if got != "" {
+ t.Errorf("want empty context, got %q", got)
+ }
+}
+
+func TestSweepContextRespectsCancellation(t *testing.T) {
+ t.Parallel()
+ size := testContextChars * 2 / 3
+ entries := []index.Entry{}
+ for i := range 5 {
+ entries = append(entries, chunkEntry(i, size))
+ }
+ ctx, cancel := context.WithCancel(context.Background())
+ cancel()
+ chat := &mockChatter{ReplyFunc: func(int, string, []llm.Message) (string, error) {
+ return "", fmt.Errorf("canceled")
+ }}
+ if _, err := sweepContext(ctx, sweepCmd(), chat, "what happened", entries, testContextChars); err == nil {
+ t.Fatal("want an error once the context is canceled")
+ }
+}
diff --git a/cmd/cmd_ask.go b/cmd/cmd_ask.go
new file mode 100644
index 0000000..60fbfa1
--- /dev/null
+++ b/cmd/cmd_ask.go
@@ -0,0 +1,205 @@
+package cmd
+
+import (
+ "bufio"
+ "errors"
+ "fmt"
+ "io"
+ "strings"
+ "time"
+
+ "github.com/spf13/cobra"
+
+ "github.com/dcadolph/midden/internal/interview"
+ "github.com/dcadolph/midden/internal/vault"
+ "github.com/dcadolph/midden/internal/weave"
+)
+
+// askedPrefix marks the line recording which question an entry answers, so a
+// question is never put twice.
+const askedPrefix = "MIDDEN-ASKED: "
+
+// askedPrefixLabel introduces the question an answer was given to, kept out of
+// the headline so the entry reads as the person's own words.
+const askedPrefixLabel = "In answer to: "
+
+// Ask options.
+var (
+ askAnswer string
+ askCount int
+ askTags []string
+ askInteractive bool
+ askList bool
+ askSkip bool
+)
+
+// askCmd puts a question the record cannot answer about itself.
+var askCmd = &cobra.Command{
+ Use: "ask",
+ Short: "Answer a question about a gap in your own record.",
+ Long: "Ask puts one question drawn from what the record proves is missing.\n\n" +
+ "Imported history reconstructs where you were and what you produced, because calendars and " +
+ "commit logs already exist. It cannot reconstruct what you thought, because nothing recorded " +
+ "that at the time, and no further import will fix it. Ask closes that gap the only way it can " +
+ "be closed: a commitment held for years stopped and nobody wrote down why, so it asks.\n\n" +
+ "Questions come from arithmetic over the record, never from a model, so nothing is ever asked " +
+ "about something that did not happen. Answers are ordinary entries and a question already " +
+ "answered is not asked again.",
+ RunE: runAsk,
+}
+
+func init() {
+ askCmd.Flags().StringVarP(&askAnswer, "answer", "a", "", "Answer the next question and file it as an entry.")
+ askCmd.Flags().IntVarP(&askCount, "count", "n", 1, "Number of questions to show.")
+ askCmd.Flags().StringSliceVarP(&askTags, "tag", "t", []string{"answer"}, "Tags to attach to the answer.")
+ askCmd.Flags().BoolVarP(&askInteractive, "interactive", "i", false, "Ask, then read the answer from stdin.")
+ askCmd.Flags().BoolVar(&askList, "list", false, "List pending questions without answering.")
+ askCmd.Flags().BoolVar(&askSkip, "skip", false,
+ "Dismiss the next question without answering it, so it is never asked again.")
+ rootCmd.AddCommand(askCmd)
+}
+
+// runAsk generates outstanding questions and either shows them or captures an
+// answer to the first.
+func runAsk(cmd *cobra.Command, _ []string) error {
+ v, err := openVault()
+ if err != nil {
+ return err
+ }
+ entries, err := entriesInRange(v, dateRange{})
+ if err != nil {
+ return errors.Join(ErrVault, fmt.Errorf("scan vault: %w", err))
+ }
+ questions := pendingQuestions(entries)
+ if len(questions) == 0 {
+ fmt.Fprintln(cmd.OutOrStdout(), "Nothing to ask: the record has no unexplained gaps yet.")
+ return nil
+ }
+
+ if askSkip {
+ q := questions[0]
+ // A dismissal is recorded the same way an answer is, because a queue that
+ // keeps returning a question the person has already rejected trains them
+ // to stop reading it at all.
+ entry := vault.Entry{
+ Time: time.Now(),
+ Tags: entryTags([]string{"skipped"}),
+ Body: fmt.Sprintf("Not worth recording.\n%s%s\n%s%s",
+ askedPrefixLabel, q.Prompt, askedPrefix, q.ID),
+ }
+ if err := v.Append(entry); err != nil {
+ return errors.Join(ErrVault, fmt.Errorf("record dismissal: %w", err))
+ }
+ fmt.Fprintf(cmd.OutOrStdout(), "Dismissed: %s\n", q.Prompt)
+ fmt.Fprintf(cmd.ErrOrStderr(), "%d question(s) still outstanding.\n", len(questions)-1)
+ return nil
+ }
+
+ if askList || (askAnswer == "" && !askInteractive) {
+ writeQuestions(cmd.OutOrStdout(), questions, askCount)
+ return nil
+ }
+
+ q := questions[0]
+ answer := askAnswer
+ if answer == "" {
+ writeQuestion(cmd.ErrOrStderr(), q)
+ answer, err = readAnswer(cmd.InOrStdin())
+ if err != nil {
+ return err
+ }
+ }
+ if strings.TrimSpace(answer) == "" {
+ return errors.Join(ErrNotFound, errors.New("no answer given"))
+ }
+ entry := vault.Entry{
+ Time: time.Now(),
+ Tags: entryTags(askTags),
+ // The answer leads and the question trails as attribution. An entry's
+ // headline is what every other command shows and what weave groups on, so
+ // putting the prompt first would make the record display midden's
+ // questions back instead of the person's own words.
+ Body: fmt.Sprintf("%s\n\n%s%s\n%s%s",
+ strings.TrimSpace(answer), askedPrefixLabel, q.Prompt, askedPrefix, q.ID),
+ }
+ if err := v.Append(entry); err != nil {
+ return errors.Join(ErrVault, fmt.Errorf("append answer: %w", err))
+ }
+ fmt.Fprintf(cmd.OutOrStdout(), "Answered: %s\n", q.Prompt)
+ fmt.Fprintf(cmd.ErrOrStderr(), "%d question(s) still outstanding.\n", len(questions)-1)
+ return nil
+}
+
+// pendingQuestions derives the outstanding questions from the record.
+func pendingQuestions(entries []vault.Entry) []interview.Question {
+ now := time.Now()
+ answered := map[string]bool{}
+ for _, e := range entries {
+ if id := bodyMarker(e.Body, askedPrefix); id != "" {
+ answered[id] = true
+ }
+ }
+ // Answers are prose rather than events, so they must not themselves become
+ // threads the interviewer then asks about.
+ source := make([]vault.Entry, 0, len(entries))
+ for _, e := range entries {
+ if bodyMarker(e.Body, askedPrefix) == "" {
+ source = append(source, e)
+ }
+ }
+ threads := weave.Threads(source, weave.DefaultOptions(now))
+ crossings := weave.Overlaps(source, []string{"calendar", "git"}, now)
+ if len(crossings) > 3 {
+ crossings = crossings[:3]
+ }
+ gaps := weave.GapsBySource(source, []string{"calendar", "git"}, weave.DefaultGapOptions(now))
+ return interview.Generate(threads, crossings, gaps, answered, interview.DefaultOptions(now))
+}
+
+// writeQuestions renders up to n outstanding questions.
+func writeQuestions(w io.Writer, questions []interview.Question, n int) {
+ if n < 1 {
+ n = 1
+ }
+ if n > len(questions) {
+ n = len(questions)
+ }
+ for _, q := range questions[:n] {
+ writeQuestion(w, q)
+ }
+ if remaining := len(questions) - n; remaining > 0 {
+ fmt.Fprintf(w, "%d more outstanding. Answer with: midden ask -i\n", remaining)
+ }
+}
+
+// writeQuestion renders one question with the evidence behind it.
+func writeQuestion(w io.Writer, q interview.Question) {
+ fmt.Fprintf(w, "\n%s\n", q.Prompt)
+ for line := range strings.SplitSeq(strings.TrimSpace(q.Context), "\n") {
+ fmt.Fprintf(w, " %s\n", line)
+ }
+ fmt.Fprintln(w)
+}
+
+// readAnswer reads a free-text answer, ending at a blank line or end of input,
+// so a spoken-length reply can run to several lines without ceremony.
+func readAnswer(r io.Reader) (string, error) {
+ fmt.Print("> ")
+ scanner := bufio.NewScanner(r)
+ scanner.Buffer(make([]byte, 64*1024), 1024*1024)
+ var lines []string
+ for scanner.Scan() {
+ line := scanner.Text()
+ if strings.TrimSpace(line) == "" && len(lines) > 0 {
+ break
+ }
+ if strings.TrimSpace(line) == "" {
+ continue
+ }
+ lines = append(lines, line)
+ }
+ if err := scanner.Err(); err != nil {
+ return "", fmt.Errorf("read answer: %w", err)
+ }
+ return strings.Join(lines, "\n"), nil
+}
diff --git a/cmd/cmd_ask_test.go b/cmd/cmd_ask_test.go
new file mode 100644
index 0000000..7e93b4a
--- /dev/null
+++ b/cmd/cmd_ask_test.go
@@ -0,0 +1,161 @@
+package cmd
+
+import (
+ "strings"
+ "testing"
+ "time"
+
+ "github.com/google/go-cmp/cmp"
+
+ "github.com/dcadolph/midden/internal/vault"
+)
+
+// askVault returns a vault holding a weekly class that ran for months and then
+// stopped, which is the shape ask is built to notice.
+func askVault(t *testing.T) *vault.Vault {
+ t.Helper()
+ v, err := vault.Open(t.TempDir())
+ if err != nil {
+ t.Fatalf("vault.Open: %v", err)
+ }
+ start := time.Now().AddDate(-2, 0, 0)
+ entries := make([]vault.Entry, 0, 40)
+ for i := range 40 {
+ entries = append(entries, vault.Entry{
+ Time: start.AddDate(0, 0, 7*i),
+ Tags: []string{"calendar"},
+ Body: "William- martial arts",
+ })
+ }
+ if err := v.AppendAll(entries); err != nil {
+ t.Fatalf("AppendAll: %v", err)
+ }
+ return v
+}
+
+// entriesOf reads every entry in the vault.
+func entriesOf(t *testing.T, v *vault.Vault) []vault.Entry {
+ t.Helper()
+ out, err := entriesInRange(v, dateRange{})
+ if err != nil {
+ t.Fatalf("entriesInRange: %v", err)
+ }
+ return out
+}
+
+func TestPendingQuestionsFindsAnEnding(t *testing.T) {
+ t.Parallel()
+ got := pendingQuestions(entriesOf(t, askVault(t)))
+ if len(got) == 0 {
+ t.Fatal("want a question about the class that stopped")
+ }
+ if !strings.Contains(got[0].Prompt, "martial arts") {
+ t.Errorf("want the ended thread asked about, got %q", got[0].Prompt)
+ }
+}
+
+func TestAnsweredQuestionsAreRetired(t *testing.T) {
+ t.Parallel()
+ v := askVault(t)
+ first := pendingQuestions(entriesOf(t, v))
+ if len(first) == 0 {
+ t.Fatal("want a question to answer")
+ }
+ // An answer entry carries the marker, which is what retires the question.
+ err := v.Append(vault.Entry{
+ Time: time.Now(),
+ Tags: []string{"answer"},
+ Body: "He switched to baseball.\n\n" + askedPrefixLabel + first[0].Prompt + "\n" + askedPrefix + first[0].ID,
+ })
+ if err != nil {
+ t.Fatalf("Append: %v", err)
+ }
+ for _, q := range pendingQuestions(entriesOf(t, v)) {
+ if q.ID == first[0].ID {
+ t.Error("want an answered question never asked again")
+ }
+ }
+}
+
+func TestDismissedQuestionsAreAlsoRetired(t *testing.T) {
+ t.Parallel()
+ v := askVault(t)
+ first := pendingQuestions(entriesOf(t, v))
+ if len(first) == 0 {
+ t.Fatal("want a question to dismiss")
+ }
+ // A queue that keeps returning a rejected question trains the person to stop
+ // reading it, so a dismissal has to stick exactly like an answer.
+ err := v.Append(vault.Entry{
+ Time: time.Now(),
+ Tags: []string{"skipped"},
+ Body: "Not worth recording.\n" + askedPrefixLabel + first[0].Prompt + "\n" + askedPrefix + first[0].ID,
+ })
+ if err != nil {
+ t.Fatalf("Append: %v", err)
+ }
+ for _, q := range pendingQuestions(entriesOf(t, v)) {
+ if q.ID == first[0].ID {
+ t.Error("want a dismissed question never asked again")
+ }
+ }
+}
+
+func TestAnswersDoNotBecomeQuestionsThemselves(t *testing.T) {
+ t.Parallel()
+ v := askVault(t)
+ q := pendingQuestions(entriesOf(t, v))
+ // Enough answers to form a thread on their own if they were not excluded.
+ for i := range 8 {
+ err := v.Append(vault.Entry{
+ Time: time.Now().AddDate(0, 0, -i*7),
+ Tags: []string{"answer"},
+ Body: "Some answer text.\n\n" + askedPrefixLabel + q[0].Prompt + "\n" + askedPrefix + "fake-" + string(rune('a'+i)),
+ })
+ if err != nil {
+ t.Fatalf("Append: %v", err)
+ }
+ }
+ // The interviewer asking about its own past questions would be a loop that
+ // never closes.
+ for _, got := range pendingQuestions(entriesOf(t, v)) {
+ if strings.Contains(got.Prompt, "Some answer text") || strings.Contains(got.Prompt, "In answer to") {
+ t.Errorf("want answers excluded from question generation, got %q", got.Prompt)
+ }
+ }
+}
+
+func TestBodyMarkerReadsTheAskedID(t *testing.T) {
+ t.Parallel()
+ body := "He switched to baseball.\n\n" + askedPrefixLabel + "William- Martial arts stopped. What happened?\n" + askedPrefix + "ended-191c9382"
+ if diff := cmp.Diff("ended-191c9382", bodyMarker(body, askedPrefix)); diff != "" {
+ t.Errorf("mismatch (-want +got):\n%s", diff)
+ }
+ // The headline is the person's own words, which is what every read command
+ // shows and what weave groups on.
+ if diff := cmp.Diff("He switched to baseball.", firstLine(body)); diff != "" {
+ t.Errorf("headline mismatch (-want +got):\n%s", diff)
+ }
+}
+
+func TestReadAnswerAcceptsSeveralLines(t *testing.T) {
+ t.Parallel()
+ got, err := readAnswer(strings.NewReader("first line\nsecond line\n\nignored after the blank\n"))
+ if err != nil {
+ t.Fatalf("readAnswer: %v", err)
+ }
+ if diff := cmp.Diff("first line\nsecond line", got); diff != "" {
+ t.Errorf("mismatch (-want +got):\n%s", diff)
+ }
+}
+
+func TestReadAnswerHandlesEmptyInput(t *testing.T) {
+ t.Parallel()
+ got, err := readAnswer(strings.NewReader(""))
+ if err != nil {
+ t.Fatalf("readAnswer: %v", err)
+ }
+ if got != "" {
+ t.Errorf("want an empty answer, got %q", got)
+ }
+}
diff --git a/cmd/cmd_chat.go b/cmd/cmd_chat.go
index ddd8801..88f80fb 100644
--- a/cmd/cmd_chat.go
+++ b/cmd/cmd_chat.go
@@ -9,28 +9,83 @@ import (
"github.com/spf13/cobra"
+ "github.com/dcadolph/midden/index"
"github.com/dcadolph/midden/llm"
)
// chatTopK caps the number of entries fed to the chat model as context.
var chatTopK int
+// chatSince and chatUntil bound the entries the answer may draw on.
+var (
+ chatSince string
+ chatUntil string
+)
+
+// chatSweep answers from every entry in scope rather than the closest matches.
+var chatSweep bool
+
+// chatContextChars caps how much entry text is sent to the model in one call.
+var chatContextChars int
+
+// Timeouts for the chat call. Current models reason before answering, so even a
+// single reply can take minutes over a large record; a sweep may summarize a
+// decade in chunks and is budgeted against that rather than against one reply.
+const (
+ chatTimeout = 5 * time.Minute
+ chatSweepTimeout = 30 * time.Minute
+)
+
+// digestTopTags caps the tag histogram supplied with every answer.
+const digestTopTags = 25
+
+// chatSystem frames what the model is reading. The two halves of the context
+// carry different authority: the summary counts every entry in scope, while the
+// entries are the quotable text and may be a sample, so the model is told not
+// to read absence from the sample as absence from the record.
+const chatSystem = "You are reading a personal journal and calendar archive. " +
+ "Answer only from the record supplied below. " +
+ "The corpus summary is a complete count over every indexed entry in scope: use it for questions " +
+ "about the shape of the record as a whole, such as which periods, people, places, or themes recur. " +
+ "The entries after it are the specific text you may quote, and they may be a sample rather than the " +
+ "whole record, so never conclude that something did not happen merely because it is absent from them. " +
+ "Quote the date of any entry you cite. Say plainly when the record does not answer the question. " +
+ "Do not invent facts."
+
// chatCmd answers a question using recalled entries as context.
var chatCmd = &cobra.Command{
Use: "chat [question...]",
Short: "Answer a question using semantic recall plus a chat model.",
- Args: cobra.MinimumNArgs(1),
- RunE: runChat,
+ Long: "Chat answers a question from the indexed vault.\n\n" +
+ "By default it retrieves the entries closest to the question. Questions about a period of time " +
+ "are answered better by scoping the range with --since and --until, which sweeps every entry in " +
+ "that range instead of ranking them, and questions about the record as a whole are answered " +
+ "better with --sweep. Every answer also receives a summary counted over the whole vault in scope.",
+ Args: cobra.MinimumNArgs(1),
+ RunE: runChat,
}
func init() {
chatCmd.Flags().IntVarP(&chatTopK, "top", "k", 8, "Number of entries to include as context.")
+ chatCmd.Flags().StringVar(&chatSince, "since", "", "Only consider entries on or after this date.")
+ chatCmd.Flags().StringVar(&chatUntil, "until", "", "Only consider entries on or before this date.")
+ chatCmd.Flags().BoolVar(&chatSweep, "sweep", false,
+ "Answer from every entry in scope instead of the closest matches, summarizing in chunks when the range is large.")
+ chatCmd.Flags().IntVar(&chatContextChars, "context-chars", defaultContextChars,
+ "Characters of entry text to send in one call. Lower this for a small local model, which is slow on large prompts.")
rootCmd.AddCommand(chatCmd)
}
-// runChat embeds the question, recalls relevant entries, and asks the chat model
-// to answer using only those entries as evidence.
+// runChat embeds the question, gathers the entries in scope, and asks the chat
+// model to answer using only those entries and the corpus summary as evidence.
func runChat(cmd *cobra.Command, args []string) error {
+ if chatContextChars < 1 {
+ return fmt.Errorf("--context-chars must be at least 1")
+ }
+ span, err := resolveDateRange(chatSince, chatUntil)
+ if err != nil {
+ return err
+ }
v, err := openVault()
if err != nil {
return err
@@ -44,29 +99,54 @@ func runChat(cmd *cobra.Command, args []string) error {
if err != nil {
return errors.Join(ErrLLM, fmt.Errorf("pick chatter: %w", err))
}
- matches := rc.Index.Search(rc.Query, chatTopK)
- var contextBuf strings.Builder
- for _, m := range matches {
- contextBuf.WriteString("--- ")
- contextBuf.WriteString(m.Entry.Time.Format(layoutDateTime))
- if len(m.Entry.Tags) > 0 {
- contextBuf.WriteString(" [")
- contextBuf.WriteString(strings.Join(m.Entry.Tags, ", "))
- contextBuf.WriteString("]")
- }
- contextBuf.WriteString("\n")
- contextBuf.WriteString(m.Entry.Body)
- contextBuf.WriteString("\n\n")
+
+ // A bounded range is a question about a period, which a ranked head cannot
+ // answer, so scoping the range implies sweeping it.
+ sweep := chatSweep || span.Bounded()
+ timeout := chatTimeout
+ if sweep {
+ timeout = chatSweepTimeout
}
- system := "You are a personal journal assistant. Answer the user's question using only the journal entries supplied below. Quote the date of any entry you cite. If the entries do not contain the answer, say so plainly. Do not invent facts."
- user := "Question: " + question + "\n\nRelevant entries:\n" + contextBuf.String()
- ctx, cancel := context.WithTimeout(context.Background(), 90*time.Second)
+ ctx, cancel := context.WithTimeout(context.Background(), timeout)
defer cancel()
- reply, err := chat.Reply(ctx, system, []llm.Message{{Role: "user", Content: user}})
+
+ var (
+ body string
+ count int
+ )
+ if sweep {
+ entries := rc.Index.InRange(span.From, span.To)
+ count = len(entries)
+ body, err = sweepContext(ctx, cmd, chat, question, entries, chatContextChars)
+ if err != nil {
+ return errors.Join(ErrLLM, err)
+ }
+ } else {
+ matches := rc.Index.SearchRange(rc.Query, chatTopK, span.From, span.To)
+ count = len(matches)
+ body = renderEntries(matchedEntries(matches))
+ }
+
+ digest := rc.Index.Digest(span.From, span.To, digestTopTags)
+ if digest.Entries == 0 {
+ return errors.Join(ErrNotFound, fmt.Errorf("no indexed entries in range (%s)", span.Label()))
+ }
+ user := "Question: " + question + "\n\n" + renderDigest(digest, span.Label()) + "\nRecord:\n" + body
+ reply, err := chat.Reply(ctx, chatSystem, []llm.Message{{Role: "user", Content: user}})
if err != nil {
return errors.Join(ErrLLM, fmt.Errorf("chat: %w", err))
}
fmt.Fprintln(cmd.OutOrStdout(), strings.TrimSpace(reply))
- fmt.Fprintf(cmd.ErrOrStderr(), "\n(answered with %s over %d entries via %s)\n", chat.Name(), len(matches), rc.Embedder.Name())
+ fmt.Fprintf(cmd.ErrOrStderr(), "\n(answered with %s over %d of %d entries in %s via %s)\n",
+ chat.Name(), count, digest.Entries, span.Label(), rc.Embedder.Name())
return nil
}
+
+// matchedEntries strips the similarity scores from ranked matches.
+func matchedEntries(matches []index.Match) []index.Entry {
+ out := make([]index.Entry, len(matches))
+ for i, m := range matches {
+ out[i] = m.Entry
+ }
+ return out
+}
diff --git a/cmd/cmd_ingest.go b/cmd/cmd_ingest.go
index 5fe5a45..0bb710d 100644
--- a/cmd/cmd_ingest.go
+++ b/cmd/cmd_ingest.go
@@ -9,11 +9,17 @@ import (
"github.com/spf13/cobra"
- "github.com/dcadolph/midden/dateutil"
"github.com/dcadolph/midden/ics"
"github.com/dcadolph/midden/internal/vault"
)
+// uidPrefix marks the line that records which calendar occurrence an entry came
+// from, so repeated ingests recognize what is already in the vault.
+const uidPrefix = "ICS-UID: "
+
+// ingestProgressEvery is how many appended events pass between progress lines.
+const ingestProgressEvery = 250
+
// ingestCmd groups passive-ingestion subcommands.
var ingestCmd = &cobra.Command{
Use: "ingest",
@@ -24,109 +30,226 @@ var ingestCmd = &cobra.Command{
var ingestICSCmd = &cobra.Command{
Use: "ics [file]",
Short: "Append calendar events from an .ics file as entries.",
- Args: cobra.ExactArgs(1),
- RunE: runIngestICS,
+ Long: "Ingest reads a calendar export and appends one entry per event occurrence.\n\n" +
+ "The whole file is ingested by default, because importing years of calendar history is the " +
+ "point of the command; narrow it with --from and --to when you want part of it. Recurring " +
+ "series are expanded into the occurrences they actually produced, so a weekly meeting " +
+ "contributes every week it happened rather than only its first. Occurrences already in the " +
+ "vault are skipped, so ingesting the same export twice is safe.",
+ Args: cobra.ExactArgs(1),
+ RunE: runIngestICS,
}
-// ingestFrom and ingestTo bound the date range pulled out of the calendar.
+// ingestSince and ingestUntil narrow the date range pulled out of the calendar.
+// Empty means the range is derived from the file itself.
var (
- ingestFrom string
- ingestTo string
- ingestTag []string
+ ingestSince string
+ ingestUntil string
+ ingestTag []string
)
func init() {
- ingestICSCmd.Flags().StringVar(&ingestFrom, "from", "today", "Start of the date range to ingest.")
- ingestICSCmd.Flags().StringVar(&ingestTo, "to", "today", "End of the date range to ingest.")
+ ingestICSCmd.Flags().StringVar(&ingestSince, "since", "",
+ "Only ingest events on or after this date (default: the earliest event in the file).")
+ ingestICSCmd.Flags().StringVar(&ingestUntil, "until", "",
+ "Only ingest events on or before this date (default: the later of the last event in the file and today).")
ingestICSCmd.Flags().StringSliceVarP(&ingestTag, "tag", "t", []string{"calendar"}, "Tags to attach to every ingested event.")
ingestCmd.AddCommand(ingestICSCmd)
rootCmd.AddCommand(ingestCmd)
}
-// runIngestICS parses the .ics file and appends an entry per event inside the chosen range.
-// Events carrying a UID are deduplicated against entries already in the vault.
+// runIngestICS parses the .ics file, expands recurring series into occurrences,
+// and appends the ones the vault does not already hold.
func runIngestICS(cmd *cobra.Command, args []string) error {
- from, err := dateutil.Parse(ingestFrom)
+ events, skipped, err := parseICSFile(args[0])
+ if err != nil {
+ return err
+ }
+ if skipped > 0 {
+ fmt.Fprintf(cmd.ErrOrStderr(),
+ "Warning: skipped %d event(s) with a missing or unparseable DTSTART.\n", skipped)
+ }
+ span, err := ingestWindow(events, ingestSince, ingestUntil)
if err != nil {
return err
}
- to, err := dateutil.Parse(ingestTo)
+ occurrences, report := ics.Expand(events, span.From, span.To)
+ reportExpansion(cmd, report)
+
+ v, err := openVault()
if err != nil {
return err
}
- to = to.AddDate(0, 0, 1).Add(-time.Nanosecond)
- f, err := os.Open(args[0])
+ seen, err := existingOccurrences(v, span)
+ if err != nil {
+ return errors.Join(ErrVault, fmt.Errorf("scan vault for existing events: %w", err))
+ }
+ tags := entryTags(ingestTag)
+ entries := make([]vault.Entry, 0, len(occurrences))
+ dupes := 0
+ for _, e := range occurrences {
+ body := formatEvent(e)
+ if seen[occurrenceKey(e.UID, firstLine(body), e.Start)] {
+ dupes++
+ continue
+ }
+ entries = append(entries, vault.Entry{Time: e.Start, Tags: tags, Body: body})
+ if len(entries)%ingestProgressEvery == 0 {
+ fmt.Fprintf(cmd.ErrOrStderr(), "Prepared %d/%d event(s)\n", len(entries), len(occurrences))
+ }
+ }
+ if err := v.AppendAll(entries); err != nil {
+ return errors.Join(ErrVault, fmt.Errorf("append events: %w", err))
+ }
+ if dupes > 0 {
+ fmt.Fprintf(cmd.ErrOrStderr(), "Skipped %d event(s) already in the vault.\n", dupes)
+ }
+ fmt.Fprintf(cmd.OutOrStdout(), "Ingested %d event(s) into the vault (%s).\n", len(entries), span.Label())
+ return nil
+}
+
+// parseICSFile reads and parses the calendar export at the given path.
+func parseICSFile(path string) ([]ics.Event, int, error) {
+ f, err := os.Open(path) //nolint:gosec // The calendar path is the command argument the user typed.
if err != nil {
- return fmt.Errorf("open %s: %w", args[0], err)
+ return nil, 0, fmt.Errorf("open %s: %w", path, err)
}
defer func() { _ = f.Close() }()
events, skipped, err := ics.Parse(f)
if err != nil {
- return errors.Join(ErrVault, fmt.Errorf("parse ics: %w", err))
+ return nil, 0, errors.Join(ErrVault, fmt.Errorf("parse ics: %w", err))
}
- if skipped > 0 {
- fmt.Fprintf(cmd.ErrOrStderr(),
- "Warning: skipped %d event(s) with a missing or unparseable DTSTART.\n", skipped)
+ return events, skipped, nil
+}
+
+// ingestWindow resolves the range to ingest. An unset bound is derived from the
+// file so the default is the whole export rather than a single day. The far end
+// also has to be concrete because a recurring rule with no UNTIL would otherwise
+// have nothing to stop it, so it reaches at least to the end of today.
+func ingestWindow(events []ics.Event, since, until string) (dateRange, error) {
+ span, err := resolveDateRange(since, until)
+ if err != nil {
+ return dateRange{}, err
}
- v, err := openVault()
+ first, last := eventBounds(events)
+ if span.From.IsZero() && !first.IsZero() {
+ span.From = dayStart(first)
+ }
+ if span.To.IsZero() {
+ end := time.Now()
+ if last.After(end) {
+ end = last
+ }
+ span.To = dayStart(end).AddDate(0, 0, 1).Add(-time.Nanosecond)
+ }
+ return span, nil
+}
+
+// eventBounds returns the earliest and latest explicit start time in the file.
+func eventBounds(events []ics.Event) (time.Time, time.Time) {
+ var first, last time.Time
+ for _, e := range events {
+ if first.IsZero() || e.Start.Before(first) {
+ first = e.Start
+ }
+ if e.Start.After(last) {
+ last = e.Start
+ }
+ }
+ return first, last
+}
+
+// reportExpansion tells the user what the expansion could not do faithfully, so
+// a partial calendar is never presented as a complete one.
+func reportExpansion(cmd *cobra.Command, r ics.ExpandReport) {
+ w := cmd.ErrOrStderr()
+ if r.Unexpanded > 0 {
+ fmt.Fprintf(w, "Warning: %d recurring series use a rule midden does not expand; "+
+ "only their first occurrence was ingested.\n", r.Unexpanded)
+ }
+ if r.Truncated > 0 {
+ fmt.Fprintf(w, "Warning: %d recurring series were cut short during expansion and may be missing occurrences.\n",
+ r.Truncated)
+ }
+ if r.Occurrences > 0 {
+ fmt.Fprintf(w, "Expanded recurring series into %d occurrence(s); %d excluded, %d replaced by overrides.\n",
+ r.Occurrences, r.Excluded, r.Overridden)
+ }
+}
+
+// forEachEntryInRange calls fn for every entry inside the window, reading each
+// day once. Ingestion asks what the vault already holds for far more incoming
+// records than there are days to hold them, so the range is walked once here
+// rather than re-read per record.
+func forEachEntryInRange(v *vault.Vault, span dateRange, fn func(vault.Entry)) error {
+ days, err := v.ListDays()
if err != nil {
return err
}
- tags := entryTags(ingestTag)
- count := 0
- dupes := 0
- recurring := 0
- for _, e := range events {
- if e.Start.Before(from) || e.Start.After(to) {
+ for _, d := range days {
+ if !span.From.IsZero() && d.Before(dayStart(span.From)) {
continue
}
- if e.Recurs {
- recurring++
+ if !span.To.IsZero() && d.After(span.To) {
+ continue
}
- if e.UID != "" {
- dup, err := dayHasUID(v, e.Start, e.UID)
- if err != nil {
- return errors.Join(ErrVault, fmt.Errorf("dedupe event %q: %w", e.Summary, err))
- }
- if dup {
- dupes++
- continue
- }
+ entries, err := v.ReadDay(d)
+ if err != nil {
+ return err
}
- entry := vault.Entry{Time: e.Start, Tags: tags, Body: formatEvent(e)}
- if err := v.Append(entry); err != nil {
- return errors.Join(ErrVault, fmt.Errorf("append event %q: %w", e.Summary, err))
+ for _, e := range entries {
+ fn(e)
}
- count++
- }
- if recurring > 0 {
- fmt.Fprintf(cmd.ErrOrStderr(),
- "Warning: %d event(s) in range recur; recurrences are not expanded.\n", recurring)
- }
- if dupes > 0 {
- fmt.Fprintf(cmd.ErrOrStderr(), "Skipped %d duplicate event(s) already in the vault.\n", dupes)
}
- fmt.Fprintf(cmd.OutOrStdout(), "Ingested %d event(s) into the vault.\n", count)
return nil
}
-// dayHasUID reports whether any entry on the event's day already contains the UID.
-func dayHasUID(v *vault.Vault, day time.Time, uid string) (bool, error) {
- entries, err := v.ReadDay(day)
+// existingOccurrences collects the calendar occurrences already in the vault
+// across the window.
+func existingOccurrences(v *vault.Vault, span dateRange) (map[string]bool, error) {
+ seen := map[string]bool{}
+ err := forEachEntryInRange(v, span, func(e vault.Entry) {
+ seen[occurrenceKey(bodyUID(e.Body), firstLine(e.Body), e.Time)] = true
+ })
if err != nil {
- return false, fmt.Errorf("read day: %w", err)
+ return nil, err
}
- for _, entry := range entries {
- if strings.Contains(entry.Body, uid) {
- return true, nil
+ return seen, nil
+}
+
+// occurrenceKey identifies one calendar occurrence. The start time is part of
+// the key because every occurrence of a series shares the series UID, so the
+// UID alone cannot tell two of them apart. Events with no UID fall back to their
+// rendered first line, which both sides of the comparison derive the same way.
+func occurrenceKey(uid, headline string, start time.Time) string {
+ id := uid
+ if id == "" {
+ id = "line:" + headline
+ }
+ return id + "@" + start.Format("2006-01-02T15:04:05")
+}
+
+// bodyUID returns the calendar UID recorded in an entry body, or empty when the
+// entry did not come from a calendar ingest.
+func bodyUID(body string) string {
+ return bodyMarker(body, uidPrefix)
+}
+
+// bodyMarker returns the value of the given ingest marker line in an entry body,
+// or empty when the body carries no such line. The prefix must start a line, so
+// prose merely mentioning it is never mistaken for a marker.
+func bodyMarker(body, prefix string) string {
+ for line := range strings.SplitSeq(body, "\n") {
+ if after, ok := strings.CutPrefix(strings.TrimSpace(line), prefix); ok {
+ return after
}
}
- return false, nil
+ return ""
}
// formatEvent renders the event body markdown for an ingested calendar entry.
-// All-day events omit clock times, and a trailing ICS-UID line makes the entry
-// discoverable for deduplication on later runs.
+// All-day events omit clock times, and a trailing UID line makes the entry
+// recognizable on later runs.
func formatEvent(e ics.Event) string {
var b strings.Builder
b.WriteString(e.Summary)
@@ -143,7 +266,7 @@ func formatEvent(e ics.Event) string {
fmt.Fprintf(&b, "\n\n%s", e.Description)
}
if e.UID != "" {
- fmt.Fprintf(&b, "\nICS-UID: %s", e.UID)
+ fmt.Fprintf(&b, "\n%s%s", uidPrefix, e.UID)
}
return b.String()
}
diff --git a/cmd/cmd_ingest_git.go b/cmd/cmd_ingest_git.go
new file mode 100644
index 0000000..1163570
--- /dev/null
+++ b/cmd/cmd_ingest_git.go
@@ -0,0 +1,281 @@
+package cmd
+
+import (
+ "errors"
+ "fmt"
+ "path/filepath"
+ "strings"
+ "time"
+
+ "github.com/go-git/go-git/v5"
+ "github.com/go-git/go-git/v5/plumbing"
+ "github.com/go-git/go-git/v5/plumbing/object"
+ "github.com/spf13/cobra"
+
+ "github.com/dcadolph/midden/internal/util"
+ "github.com/dcadolph/midden/internal/vault"
+)
+
+// commitPrefix marks the line recording which commit an entry came from, so
+// repeated ingests recognize what the vault already holds.
+const commitPrefix = "GIT-COMMIT: "
+
+// ingestGitCmd ingests commit history from local repositories.
+var ingestGitCmd = &cobra.Command{
+ Use: "git [repo...]",
+ Short: "Append commit history from local git repositories as entries.",
+ Long: "Ingest git reads commit history and appends one entry per commit.\n\n" +
+ "Commit history is a record of what you were working on and when, which calendar exports do not " +
+ "carry. Author dates are used rather than commit dates, so rebased or cherry-picked work still " +
+ "lands on the day it was written. Merge commits are skipped unless asked for, and commits already " +
+ "in the vault are skipped, so ingesting the same repository twice is safe.",
+ Args: cobra.MinimumNArgs(1),
+ RunE: runIngestGit,
+}
+
+// Git ingestion options.
+var (
+ ingestGitSince string
+ ingestGitUntil string
+ ingestGitAuthor []string
+ ingestGitTag []string
+ ingestGitRev string
+ ingestGitMerges bool
+ ingestGitStat bool
+)
+
+func init() {
+ ingestGitCmd.Flags().StringVar(&ingestGitSince, "since", "", "Only ingest commits authored on or after this date.")
+ ingestGitCmd.Flags().StringVar(&ingestGitUntil, "until", "", "Only ingest commits authored on or before this date.")
+ ingestGitCmd.Flags().StringSliceVar(&ingestGitAuthor, "author", nil,
+ "Only ingest commits whose author name or email contains one of these values.")
+ ingestGitCmd.Flags().StringSliceVarP(&ingestGitTag, "tag", "t", []string{"git"}, "Tags to attach to every ingested commit.")
+ ingestGitCmd.Flags().StringVar(&ingestGitRev, "rev", "", "Revision to walk (default: the repository HEAD).")
+ ingestGitCmd.Flags().BoolVar(&ingestGitMerges, "merges", false, "Include merge commits.")
+ ingestGitCmd.Flags().BoolVar(&ingestGitStat, "stat", false, "Include changed-file and line counts, at the cost of a diff per commit.")
+ ingestCmd.AddCommand(ingestGitCmd)
+}
+
+// gitFilter is the set of conditions a commit must satisfy to be ingested.
+type gitFilter struct {
+ // Span bounds the author dates accepted.
+ Span dateRange
+ // Authors matches against the commit author name and email; empty accepts any.
+ Authors []string
+ // Rev is the revision to walk, or empty for HEAD.
+ Rev string
+ // Merges includes merge commits when set.
+ Merges bool
+ // Stat includes changed-file and line counts when set.
+ Stat bool
+}
+
+// runIngestGit walks each repository and appends the commits the vault does not
+// already hold.
+func runIngestGit(cmd *cobra.Command, args []string) error {
+ span, err := resolveDateRange(ingestGitSince, ingestGitUntil)
+ if err != nil {
+ return err
+ }
+ filter := gitFilter{
+ Span: span,
+ Authors: ingestGitAuthor,
+ Rev: ingestGitRev,
+ Merges: ingestGitMerges,
+ Stat: ingestGitStat,
+ }
+ v, err := openVault()
+ if err != nil {
+ return err
+ }
+ seen, err := existingCommits(v, span)
+ if err != nil {
+ return errors.Join(ErrVault, fmt.Errorf("scan vault for existing commits: %w", err))
+ }
+ tags := entryTags(ingestGitTag)
+
+ var entries []vault.Entry
+ dupes, failed := 0, 0
+ for _, repo := range args {
+ commits, err := readCommits(repo, filter)
+ if err != nil {
+ // One unreadable repository must not cost the user the other thirty-nine.
+ // An empty repository has no HEAD at all, and a backfill sweeping a source
+ // directory will meet those routinely.
+ failed++
+ fmt.Fprintf(cmd.ErrOrStderr(), "%s: skipped (%v)\n", repoName(repo), err)
+ continue
+ }
+ kept := 0
+ for _, c := range commits {
+ if seen[c.SHA] {
+ dupes++
+ continue
+ }
+ seen[c.SHA] = true
+ entries = append(entries, vault.Entry{Time: c.When, Tags: tags, Body: c.Body})
+ kept++
+ }
+ fmt.Fprintf(cmd.ErrOrStderr(), "%s: %d commit(s) in range, %d new\n", repoName(repo), len(commits), kept)
+ }
+ if failed == len(args) {
+ return errors.Join(ErrGit, fmt.Errorf("every repository failed to read (%d of %d)", failed, len(args)))
+ }
+ if err := v.AppendAll(entries); err != nil {
+ return errors.Join(ErrVault, fmt.Errorf("append commits: %w", err))
+ }
+ if dupes > 0 {
+ fmt.Fprintf(cmd.ErrOrStderr(), "Skipped %d commit(s) already in the vault.\n", dupes)
+ }
+ if failed > 0 {
+ fmt.Fprintf(cmd.ErrOrStderr(), "Skipped %d unreadable repository/repositories.\n", failed)
+ }
+ fmt.Fprintf(cmd.OutOrStdout(), "Ingested %d commit(s) into the vault (%s).\n", len(entries), span.Label())
+ return nil
+}
+
+// gitCommit is one commit rendered for the vault.
+type gitCommit struct {
+ // SHA is the full commit hash, which identifies the commit for deduplication.
+ SHA string
+ // When is the author timestamp, which is when the work was actually written.
+ When time.Time
+ // Body is the rendered entry text.
+ Body string
+}
+
+// readCommits walks the repository and returns the commits passing the filter,
+// ordered oldest first.
+func readCommits(path string, filter gitFilter) ([]gitCommit, error) {
+ repo, err := git.PlainOpen(path)
+ if err != nil {
+ return nil, fmt.Errorf("open repository: %w", err)
+ }
+ start, err := resolveRev(repo, filter.Rev)
+ if err != nil {
+ return nil, err
+ }
+ iter, err := repo.Log(&git.LogOptions{From: start, Order: git.LogOrderCommitterTime})
+ if err != nil {
+ return nil, fmt.Errorf("read log: %w", err)
+ }
+ defer iter.Close()
+ name := repoName(path)
+ var out []gitCommit
+ err = iter.ForEach(func(c *object.Commit) error {
+ when := c.Author.When.Local()
+ if !filter.accepts(c, when) {
+ return nil
+ }
+ out = append(out, gitCommit{SHA: c.Hash.String(), When: when, Body: formatCommit(c, name, filter.Stat)})
+ return nil
+ })
+ if err != nil {
+ return nil, fmt.Errorf("walk log: %w", err)
+ }
+ // The log walks newest first; the vault reads better oldest first.
+ for i, j := 0, len(out)-1; i < j; i, j = i+1, j-1 {
+ out[i], out[j] = out[j], out[i]
+ }
+ return out, nil
+}
+
+// accepts reports whether the commit passes every condition in the filter.
+func (f gitFilter) accepts(c *object.Commit, when time.Time) bool {
+ if !f.Merges && c.NumParents() > 1 {
+ return false
+ }
+ if !f.Span.From.IsZero() && when.Before(f.Span.From) {
+ return false
+ }
+ if !f.Span.To.IsZero() && when.After(f.Span.To) {
+ return false
+ }
+ if len(f.Authors) == 0 {
+ return true
+ }
+ who := c.Author.Name + " <" + c.Author.Email + ">"
+ for _, want := range f.Authors {
+ if util.ContainsFold(who, want) {
+ return true
+ }
+ }
+ return false
+}
+
+// resolveRev returns the hash to start the log walk from.
+func resolveRev(repo *git.Repository, rev string) (plumbing.Hash, error) {
+ if rev == "" {
+ head, err := repo.Head()
+ if err != nil {
+ return plumbing.ZeroHash, fmt.Errorf("resolve HEAD: %w", err)
+ }
+ return head.Hash(), nil
+ }
+ hash, err := repo.ResolveRevision(plumbing.Revision(rev))
+ if err != nil {
+ return plumbing.ZeroHash, fmt.Errorf("resolve %q: %w", rev, err)
+ }
+ return *hash, nil
+}
+
+// formatCommit renders the entry body for one commit. The subject leads so the
+// entry reads as a line of history, and the trailing hash line makes the commit
+// recognizable on later runs.
+func formatCommit(c *object.Commit, repo string, withStat bool) string {
+ message := strings.TrimSpace(c.Message)
+ subject, rest, _ := strings.Cut(message, "\n")
+ var b strings.Builder
+ b.WriteString(strings.TrimSpace(subject))
+ fmt.Fprintf(&b, "\nRepo: %s", repo)
+ fmt.Fprintf(&b, "\nAuthor: %s", c.Author.Name)
+ if withStat {
+ if summary := statSummary(c); summary != "" {
+ fmt.Fprintf(&b, "\nChanges: %s", summary)
+ }
+ }
+ if body := strings.TrimSpace(rest); body != "" {
+ fmt.Fprintf(&b, "\n\n%s", body)
+ }
+ fmt.Fprintf(&b, "\n%s%s", commitPrefix, c.Hash.String())
+ return b.String()
+}
+
+// statSummary renders the changed-file and line counts for a commit, or empty
+// when the diff cannot be computed.
+func statSummary(c *object.Commit) string {
+ stats, err := c.Stats()
+ if err != nil {
+ return ""
+ }
+ added, deleted := 0, 0
+ for _, s := range stats {
+ added += s.Addition
+ deleted += s.Deletion
+ }
+ return fmt.Sprintf("%d file(s), +%d/-%d", len(stats), added, deleted)
+}
+
+// repoName returns the directory name identifying the repository in an entry.
+func repoName(path string) string {
+ abs, err := filepath.Abs(path)
+ if err != nil {
+ return filepath.Base(path)
+ }
+ return filepath.Base(filepath.Clean(abs))
+}
+
+// existingCommits collects the commit hashes already recorded in the vault
+// across the window.
+func existingCommits(v *vault.Vault, span dateRange) (map[string]bool, error) {
+ seen := map[string]bool{}
+ err := forEachEntryInRange(v, span, func(e vault.Entry) {
+ if sha := bodyMarker(e.Body, commitPrefix); sha != "" {
+ seen[sha] = true
+ }
+ })
+ if err != nil {
+ return nil, err
+ }
+ return seen, nil
+}
diff --git a/cmd/cmd_ingest_git_test.go b/cmd/cmd_ingest_git_test.go
new file mode 100644
index 0000000..68b313f
--- /dev/null
+++ b/cmd/cmd_ingest_git_test.go
@@ -0,0 +1,251 @@
+package cmd
+
+import (
+ "fmt"
+ "os"
+ "path/filepath"
+ "strings"
+ "testing"
+ "time"
+
+ "github.com/go-git/go-git/v5"
+ "github.com/go-git/go-git/v5/plumbing"
+ "github.com/go-git/go-git/v5/plumbing/object"
+ "github.com/google/go-cmp/cmp"
+
+ "github.com/dcadolph/midden/internal/vault"
+)
+
+// commitSpec describes one commit to write into a test repository.
+type commitSpec struct {
+ // Message is the full commit message.
+ Message string
+ // Name and Email identify the author.
+ Name string
+ Email string
+ // When is the author timestamp.
+ When time.Time
+}
+
+// testRepo builds a repository holding the given commits, in order, and returns
+// its path.
+func testRepo(t *testing.T, commits ...commitSpec) string {
+ t.Helper()
+ dir := t.TempDir()
+ repo, err := git.PlainInit(dir, false)
+ if err != nil {
+ t.Fatalf("PlainInit: %v", err)
+ }
+ wt, err := repo.Worktree()
+ if err != nil {
+ t.Fatalf("Worktree: %v", err)
+ }
+ for i, c := range commits {
+ name := fmt.Sprintf("file%d.txt", i)
+ if err := os.WriteFile(filepath.Join(dir, name), []byte(c.Message), 0o600); err != nil {
+ t.Fatalf("write %s: %v", name, err)
+ }
+ if _, err := wt.Add(name); err != nil {
+ t.Fatalf("Add %s: %v", name, err)
+ }
+ _, err := wt.Commit(c.Message, &git.CommitOptions{
+ Author: &object.Signature{Name: c.Name, Email: c.Email, When: c.When},
+ })
+ if err != nil {
+ t.Fatalf("Commit %q: %v", c.Message, err)
+ }
+ }
+ return dir
+}
+
+// gitDay returns a local timestamp on the given March 2024 day.
+func gitDay(day, hour int) time.Time {
+ return time.Date(2024, time.March, day, hour, 0, 0, 0, time.Local)
+}
+
+func TestReadCommitsOrdersOldestFirst(t *testing.T) {
+ t.Parallel()
+ dir := testRepo(t,
+ commitSpec{Message: "first", Name: "Ada", Email: "ada@example.com", When: gitDay(4, 9)},
+ commitSpec{Message: "second", Name: "Ada", Email: "ada@example.com", When: gitDay(6, 11)},
+ commitSpec{Message: "third", Name: "Ada", Email: "ada@example.com", When: gitDay(8, 15)},
+ )
+ got, err := readCommits(dir, gitFilter{})
+ if err != nil {
+ t.Fatalf("readCommits: %v", err)
+ }
+ subjects := make([]string, len(got))
+ for i, c := range got {
+ subjects[i] = firstLine(c.Body)
+ }
+ if diff := cmp.Diff([]string{"first", "second", "third"}, subjects); diff != "" {
+ t.Errorf("order mismatch (-want +got):\n%s", diff)
+ }
+ // Author time is what puts a commit on the day the work was done, which is
+ // the point of ingesting history at all.
+ if !got[0].When.Equal(gitDay(4, 9)) {
+ t.Errorf("want the author timestamp %s, got %s", gitDay(4, 9), got[0].When)
+ }
+}
+
+func TestReadCommitsHonorsDateRange(t *testing.T) {
+ t.Parallel()
+ dir := testRepo(t,
+ commitSpec{Message: "before", Name: "Ada", Email: "ada@example.com", When: gitDay(1, 9)},
+ commitSpec{Message: "inside", Name: "Ada", Email: "ada@example.com", When: gitDay(10, 9)},
+ commitSpec{Message: "after", Name: "Ada", Email: "ada@example.com", When: gitDay(20, 9)},
+ )
+ span, err := resolveDateRange("2024-03-05", "2024-03-15")
+ if err != nil {
+ t.Fatalf("resolveDateRange: %v", err)
+ }
+ got, err := readCommits(dir, gitFilter{Span: span})
+ if err != nil {
+ t.Fatalf("readCommits: %v", err)
+ }
+ if len(got) != 1 || firstLine(got[0].Body) != "inside" {
+ t.Fatalf("want only the in-range commit, got %d", len(got))
+ }
+}
+
+func TestReadCommitsFiltersByAuthor(t *testing.T) {
+ t.Parallel()
+ dir := testRepo(t,
+ commitSpec{Message: "mine", Name: "Ada Lovelace", Email: "ada@example.com", When: gitDay(4, 9)},
+ commitSpec{Message: "theirs", Name: "Grace Hopper", Email: "grace@example.com", When: gitDay(5, 9)},
+ )
+ got, err := readCommits(dir, gitFilter{Authors: []string{"ada@example.com"}})
+ if err != nil {
+ t.Fatalf("readCommits: %v", err)
+ }
+ if len(got) != 1 || firstLine(got[0].Body) != "mine" {
+ t.Fatalf("want only the matching author's commit, got %d", len(got))
+ }
+}
+
+func TestReadCommitsRejectsAnUnknownRevision(t *testing.T) {
+ t.Parallel()
+ dir := testRepo(t, commitSpec{Message: "only", Name: "Ada", Email: "ada@example.com", When: gitDay(4, 9)})
+ if _, err := readCommits(dir, gitFilter{Rev: "no-such-branch"}); err == nil {
+ t.Error("want an error for an unresolvable revision")
+ }
+}
+
+func TestReadCommitsRejectsANonRepository(t *testing.T) {
+ t.Parallel()
+ if _, err := readCommits(t.TempDir(), gitFilter{}); err == nil {
+ t.Error("want an error for a directory that is not a repository")
+ }
+}
+
+func TestGitFilterAccepts(t *testing.T) {
+ t.Parallel()
+ span, err := resolveDateRange("2024-03-05", "2024-03-15")
+ if err != nil {
+ t.Fatalf("resolveDateRange: %v", err)
+ }
+ merge := []plumbing.Hash{plumbing.NewHash("a"), plumbing.NewHash("b")}
+ tests := []struct {
+ Filter gitFilter
+ Parents []plumbing.Hash
+ Name string
+ Email string
+ When time.Time
+ Want bool
+ }{{ // Test 0: A plain commit with no filter is accepted.
+ When: gitDay(10, 9), Want: true,
+ }, { // Test 1: Merge commits are noise in a personal history and are skipped by default.
+ Parents: merge, When: gitDay(10, 9), Want: false,
+ }, { // Test 2: Merge commits are kept when asked for.
+ Filter: gitFilter{Merges: true}, Parents: merge, When: gitDay(10, 9), Want: true,
+ }, { // Test 3: A commit before the range is rejected.
+ Filter: gitFilter{Span: span}, When: gitDay(1, 9), Want: false,
+ }, { // Test 4: A commit after the range is rejected.
+ Filter: gitFilter{Span: span}, When: gitDay(20, 9), Want: false,
+ }, { // Test 5: A commit inside the range is accepted.
+ Filter: gitFilter{Span: span}, When: gitDay(10, 9), Want: true,
+ }, { // Test 6: The author filter matches an email substring.
+ Filter: gitFilter{Authors: []string{"ada@"}},
+ Name: "Ada Lovelace", Email: "ada@example.com", When: gitDay(10, 9), Want: true,
+ }, { // Test 7: The author filter matches a name regardless of case.
+ Filter: gitFilter{Authors: []string{"lovelace"}},
+ Name: "Ada Lovelace", Email: "ada@example.com", When: gitDay(10, 9), Want: true,
+ }, { // Test 8: A commit by someone else is rejected.
+ Filter: gitFilter{Authors: []string{"ada@"}},
+ Name: "Grace Hopper", Email: "grace@example.com", When: gitDay(10, 9), Want: false,
+ }, { // Test 9: Any one of several authors matching is enough.
+ Filter: gitFilter{Authors: []string{"ada@", "grace@"}},
+ Name: "Grace Hopper", Email: "grace@example.com", When: gitDay(10, 9), Want: true,
+ }}
+ for testNum, test := range tests {
+ t.Run(fmt.Sprintf("test %d", testNum), func(t *testing.T) {
+ t.Parallel()
+ c := &object.Commit{
+ ParentHashes: test.Parents,
+ Author: object.Signature{Name: test.Name, Email: test.Email, When: test.When},
+ }
+ if diff := cmp.Diff(test.Want, test.Filter.accepts(c, test.When)); diff != "" {
+ t.Errorf("mismatch (-want +got):\n%s", diff)
+ }
+ })
+ }
+}
+
+func TestFormatCommitRoundTripsThroughBodyMarker(t *testing.T) {
+ t.Parallel()
+ dir := testRepo(t, commitSpec{
+ Message: "Add retry logic\n\nThe provider rate limits under load.",
+ Name: "Ada", Email: "ada@example.com", When: gitDay(4, 9),
+ })
+ got, err := readCommits(dir, gitFilter{})
+ if err != nil {
+ t.Fatalf("readCommits: %v", err)
+ }
+ body := got[0].Body
+ if firstLine(body) != "Add retry logic" {
+ t.Errorf("want the subject to lead the entry, got %q", firstLine(body))
+ }
+ for _, want := range []string{"Repo: " + filepath.Base(dir), "Author: Ada", "The provider rate limits under load."} {
+ if !strings.Contains(body, want) {
+ t.Errorf("entry missing %q:\n%s", want, body)
+ }
+ }
+ // The hash a run writes and the hash a later run reads back have to agree, or
+ // every re-ingest duplicates the whole history.
+ if diff := cmp.Diff(got[0].SHA, bodyMarker(body, commitPrefix)); diff != "" {
+ t.Errorf("mismatch (-want +got):\n%s", diff)
+ }
+}
+
+func TestBodyMarkerIgnoresProse(t *testing.T) {
+ t.Parallel()
+ // A marker only counts when it starts a line, so a journal entry that talks
+ // about the format is never mistaken for an ingested record.
+ body := "Wrote about how GIT-COMMIT: lines work in the vault today."
+ if got := bodyMarker(body, commitPrefix); got != "" {
+ t.Errorf("want no marker from prose, got %q", got)
+ }
+}
+
+func TestExistingCommitsFindsIngestedHashes(t *testing.T) {
+ t.Parallel()
+ v, _ := reindexVault(t)
+ sha := "1a2b3c4d5e6f7a8b9c0d1e2f3a4b5c6d7e8f9a0b"
+ err := v.AppendAll([]vault.Entry{
+ {Time: gitDay(4, 9), Body: "Add retry logic\nRepo: midden\n" + commitPrefix + sha},
+ {Time: gitDay(5, 9), Body: "A handwritten entry with no marker."},
+ })
+ if err != nil {
+ t.Fatalf("AppendAll: %v", err)
+ }
+ seen, err := existingCommits(v, dateRange{})
+ if err != nil {
+ t.Fatalf("existingCommits: %v", err)
+ }
+ if !seen[sha] {
+ t.Errorf("want the ingested commit recognized, got %v", seen)
+ }
+ if len(seen) != 1 {
+ t.Errorf("want only marker-bearing entries counted, got %d", len(seen))
+ }
+}
diff --git a/cmd/cmd_ingest_test.go b/cmd/cmd_ingest_test.go
new file mode 100644
index 0000000..0738a3f
--- /dev/null
+++ b/cmd/cmd_ingest_test.go
@@ -0,0 +1,141 @@
+package cmd
+
+import (
+ "fmt"
+ "testing"
+ "time"
+
+ "github.com/google/go-cmp/cmp"
+
+ "github.com/dcadolph/midden/ics"
+)
+
+func TestIngestWindowDefaultsToTheWholeFile(t *testing.T) {
+ t.Parallel()
+ events := []ics.Event{
+ {Start: time.Date(2018, time.April, 9, 10, 0, 0, 0, time.Local)},
+ {Start: time.Date(2021, time.July, 2, 8, 0, 0, 0, time.Local)},
+ }
+ span, err := ingestWindow(events, "", "")
+ if err != nil {
+ t.Fatalf("ingestWindow: %v", err)
+ }
+ // Backfilling a decade of calendar is the point of the command, so an
+ // unnarrowed run must not collapse to a single day.
+ want := time.Date(2018, time.April, 9, 0, 0, 0, 0, time.Local)
+ if !span.From.Equal(want) {
+ t.Errorf("want the range to open at the earliest event %s, got %s", want, span.From)
+ }
+ if span.To.Before(time.Now()) {
+ t.Errorf("want the range to reach at least today, got %s", span.To)
+ }
+}
+
+func TestIngestWindowReachesPastTheLastEventForOpenSeries(t *testing.T) {
+ t.Parallel()
+ // Everything in the file is historical, but an unbounded weekly series still
+ // runs to today, so the window has to as well.
+ events := []ics.Event{{Start: time.Date(2019, time.January, 7, 9, 0, 0, 0, time.Local)}}
+ span, err := ingestWindow(events, "", "")
+ if err != nil {
+ t.Fatalf("ingestWindow: %v", err)
+ }
+ if span.To.Before(dayStart(time.Now())) {
+ t.Errorf("want the range to reach today, got %s", span.To)
+ }
+}
+
+func TestIngestWindowHonorsExplicitBounds(t *testing.T) {
+ t.Parallel()
+ events := []ics.Event{{Start: time.Date(2018, time.April, 9, 10, 0, 0, 0, time.Local)}}
+ span, err := ingestWindow(events, "2020-01-01", "2020-12-31")
+ if err != nil {
+ t.Fatalf("ingestWindow: %v", err)
+ }
+ if diff := cmp.Diff("2020-01-01 to 2020-12-31", span.Label()); diff != "" {
+ t.Errorf("mismatch (-want +got):\n%s", diff)
+ }
+}
+
+func TestIngestWindowRejectsBadBounds(t *testing.T) {
+ t.Parallel()
+ if _, err := ingestWindow(nil, "not-a-date", ""); err == nil {
+ t.Error("want an error for an unparseable --from")
+ }
+}
+
+func TestOccurrenceKeySeparatesOccurrencesOfOneSeries(t *testing.T) {
+ t.Parallel()
+ uid := "standup@example"
+ first := occurrenceKey(uid, "Standup", time.Date(2024, time.March, 4, 9, 0, 0, 0, time.Local))
+ second := occurrenceKey(uid, "Standup", time.Date(2024, time.March, 11, 9, 0, 0, 0, time.Local))
+ // Every occurrence of a series carries the same UID, so a key built from the
+ // UID alone would collapse a weekly meeting into one entry.
+ if first == second {
+ t.Errorf("want distinct keys for two occurrences, both were %q", first)
+ }
+ repeat := occurrenceKey(uid, "Standup", time.Date(2024, time.March, 4, 9, 0, 0, 0, time.Local))
+ if first != repeat {
+ t.Errorf("want a stable key for the same occurrence, got %q then %q", first, repeat)
+ }
+}
+
+func TestOccurrenceKeyFallsBackToTheHeadline(t *testing.T) {
+ t.Parallel()
+ start := time.Date(2024, time.March, 4, 9, 0, 0, 0, time.Local)
+ a := occurrenceKey("", "Lunch with Sam", start)
+ b := occurrenceKey("", "Dentist", start)
+ if a == b {
+ t.Error("want different keys for different untitled events at the same time")
+ }
+ if a != occurrenceKey("", "Lunch with Sam", start) {
+ t.Error("want a stable key for the same untitled event")
+ }
+}
+
+func TestBodyUID(t *testing.T) {
+ t.Parallel()
+ tests := []struct {
+ Body string
+ Want string
+ }{{ // Test 0: The UID line is read out of a formatted calendar entry.
+ Body: "Standup (09:00 to 09:30)\nICS-UID: abc123@google.com",
+ Want: "abc123@google.com",
+ }, { // Test 1: A UID after a description is still found.
+ Body: "Trip\nLocation: Denver\n\nPacking list\nICS-UID: xyz",
+ Want: "xyz",
+ }, { // Test 2: An entry written by hand has no UID.
+ Body: "Thought about the roadmap today.",
+ Want: "",
+ }, { // Test 3: A body mentioning the prefix mid-line does not produce a UID.
+ Body: "Wrote about the ICS-UID: format in my notes today.",
+ Want: "",
+ }}
+ for testNum, test := range tests {
+ t.Run(fmt.Sprintf("test %d", testNum), func(t *testing.T) {
+ t.Parallel()
+ if diff := cmp.Diff(test.Want, bodyUID(test.Body)); diff != "" {
+ t.Errorf("mismatch (-want +got):\n%s", diff)
+ }
+ })
+ }
+}
+
+func TestFormatEventRoundTripsThroughBodyUID(t *testing.T) {
+ t.Parallel()
+ e := ics.Event{
+ UID: "standup@example",
+ Summary: "Standup",
+ Start: time.Date(2024, time.March, 4, 9, 0, 0, 0, time.Local),
+ End: time.Date(2024, time.March, 4, 9, 30, 0, 0, time.Local),
+ Location: "Zoom",
+ }
+ body := formatEvent(e)
+ // The key an ingest writes and the key a later ingest reads back have to
+ // agree, or every re-ingest duplicates the whole calendar.
+ written := occurrenceKey(e.UID, firstLine(body), e.Start)
+ read := occurrenceKey(bodyUID(body), firstLine(body), e.Start)
+ if diff := cmp.Diff(written, read); diff != "" {
+ t.Errorf("mismatch (-want +got):\n%s", diff)
+ }
+}
diff --git a/cmd/cmd_last.go b/cmd/cmd_last.go
index d4cc84a..2da1a6b 100644
--- a/cmd/cmd_last.go
+++ b/cmd/cmd_last.go
@@ -1,6 +1,8 @@
package cmd
import (
+ "time"
+
"errors"
"fmt"
@@ -17,8 +19,13 @@ var lastCmd = &cobra.Command{
RunE: runLast,
}
+// lastFuture includes entries dated after now, which an imported calendar holds.
+var lastFuture bool
+
func init() {
lastCmd.Flags().IntVarP(&lastCount, "count", "n", 1, "Number of recent entries to show.")
+ lastCmd.Flags().BoolVar(&lastFuture, "future", false,
+ "Include entries dated after now, such as calendar appointments that have not happened yet.")
rootCmd.AddCommand(lastCmd)
}
@@ -28,7 +35,11 @@ func runLast(cmd *cobra.Command, _ []string) error {
if err != nil {
return err
}
- entries, err := v.Recent(lastCount)
+ cutoff := time.Time{}
+ if !lastFuture {
+ cutoff = time.Now()
+ }
+ entries, err := v.RecentBefore(lastCount, cutoff)
if err != nil {
return errors.Join(ErrVault, fmt.Errorf("read recent: %w", err))
}
diff --git a/cmd/cmd_recall.go b/cmd/cmd_recall.go
index eed0df9..b0f548b 100644
--- a/cmd/cmd_recall.go
+++ b/cmd/cmd_recall.go
@@ -12,6 +12,12 @@ import (
// recallTopK caps the number of entries returned.
var recallTopK int
+// recallSince and recallUntil bound the entries the search may return.
+var (
+ recallSince string
+ recallUntil string
+)
+
// recallCmd performs a semantic search over the indexed entries.
var recallCmd = &cobra.Command{
Use: "recall [query...]",
@@ -22,11 +28,18 @@ var recallCmd = &cobra.Command{
func init() {
recallCmd.Flags().IntVarP(&recallTopK, "top", "k", 5, "Number of entries to return.")
+ recallCmd.Flags().StringVar(&recallSince, "since", "", "Only consider entries on or after this date.")
+ recallCmd.Flags().StringVar(&recallUntil, "until", "", "Only consider entries on or before this date.")
rootCmd.AddCommand(recallCmd)
}
-// runRecall embeds the query, searches the index, and prints the top entries.
+// runRecall embeds the query, searches the index within the requested date
+// range, and prints the top entries.
func runRecall(cmd *cobra.Command, args []string) error {
+ span, err := resolveDateRange(recallSince, recallUntil)
+ if err != nil {
+ return err
+ }
v, err := openVault()
if err != nil {
return err
@@ -35,7 +48,7 @@ func runRecall(cmd *cobra.Command, args []string) error {
if err != nil {
return err
}
- matches := rc.Index.Search(rc.Query, recallTopK)
+ matches := rc.Index.SearchRange(rc.Query, recallTopK, span.From, span.To)
out := make([]vault.Entry, len(matches))
for i, m := range matches {
out[i] = vault.Entry{Time: m.Entry.Time, Tags: m.Entry.Tags, Body: m.Entry.Body}
diff --git a/cmd/cmd_recent.go b/cmd/cmd_recent.go
index 7f930dc..f840c23 100644
--- a/cmd/cmd_recent.go
+++ b/cmd/cmd_recent.go
@@ -1,6 +1,8 @@
package cmd
import (
+ "time"
+
"errors"
"fmt"
@@ -17,8 +19,13 @@ var recentCmd = &cobra.Command{
RunE: runRecent,
}
+// recentFuture includes entries dated after now, which an imported calendar holds.
+var recentFuture bool
+
func init() {
recentCmd.Flags().IntVarP(&recentCount, "count", "n", 10, "Number of entries to show.")
+ recentCmd.Flags().BoolVar(&recentFuture, "future", false,
+ "Include entries dated after now, such as calendar appointments that have not happened yet.")
rootCmd.AddCommand(recentCmd)
}
@@ -28,7 +35,11 @@ func runRecent(cmd *cobra.Command, _ []string) error {
if err != nil {
return err
}
- entries, err := v.Recent(recentCount)
+ cutoff := time.Time{}
+ if !recentFuture {
+ cutoff = time.Now()
+ }
+ entries, err := v.RecentBefore(recentCount, cutoff)
if err != nil {
return errors.Join(ErrVault, fmt.Errorf("read recent: %w", err))
}
diff --git a/cmd/cmd_reindex.go b/cmd/cmd_reindex.go
index 627b143..de24e84 100644
--- a/cmd/cmd_reindex.go
+++ b/cmd/cmd_reindex.go
@@ -13,26 +13,61 @@ import (
"github.com/dcadolph/midden/llm"
)
-// reindexBatch sets the number of entries embedded per provider call.
-var reindexBatch int
+// Reindex tuning.
+var (
+ // reindexBatch sets the number of entries embedded per provider call.
+ reindexBatch int
+ // reindexTimeout bounds a single provider call rather than the whole run,
+ // because a backfilled vault takes far longer to embed than any one batch.
+ reindexTimeout time.Duration
+ // reindexFull re-embeds every entry instead of reusing cached vectors.
+ reindexFull bool
+)
+
+// reindexConfig is the tuning one rebuild runs under. It is passed rather than
+// read from the flag variables so the embedding logic has no global state.
+type reindexConfig struct {
+ // Batch is the maximum number of entries embedded per provider call.
+ Batch int
+ // Timeout bounds a single provider call.
+ Timeout time.Duration
+ // Full re-embeds every entry instead of reusing cached vectors.
+ Full bool
+}
+
+// reindexCheckpointEvery is how many batches complete between index writes.
+// Checkpointing means an interrupted rebuild keeps the embeddings it already
+// paid for, and the next run picks up from there rather than starting over.
+const reindexCheckpointEvery = 20
// reindexCmd rebuilds the vector index that backs midden recall.
var reindexCmd = &cobra.Command{
Use: "reindex",
Short: "Rebuild the vector index used by midden recall.",
- RunE: runReindex,
+ Long: "Reindex embeds every entry in the vault so recall and chat can search it.\n\n" +
+ "Entries whose text is already indexed reuse their existing vector, so a rebuild after adding " +
+ "a day costs one provider call rather than re-embedding the whole vault. Progress is written to " +
+ "the index periodically, so an interrupted rebuild resumes instead of starting over. " +
+ "Use --full to re-embed everything, which is needed only after changing embedding provider.",
+ RunE: runReindex,
}
func init() {
reindexCmd.Flags().IntVar(&reindexBatch, "batch", 64, "Maximum entries embedded per provider call.")
+ reindexCmd.Flags().DurationVar(&reindexTimeout, "batch-timeout", 2*time.Minute, "Time limit for a single provider call.")
+ reindexCmd.Flags().BoolVar(&reindexFull, "full", false, "Re-embed every entry instead of reusing cached vectors.")
rootCmd.AddCommand(reindexCmd)
}
-// runReindex walks every entry, embeds the body, and writes the index to disk.
+// runReindex walks every entry, embeds the bodies it has no vector for, and
+// writes the index to disk.
func runReindex(cmd *cobra.Command, _ []string) error {
if reindexBatch < 1 {
return fmt.Errorf("--batch must be at least 1")
}
+ if reindexTimeout <= 0 {
+ return fmt.Errorf("--batch-timeout must be positive")
+ }
v, err := openVault()
if err != nil {
return err
@@ -45,40 +80,137 @@ func runReindex(cmd *cobra.Command, _ []string) error {
if err != nil {
return errors.Join(ErrVault, fmt.Errorf("scan vault: %w", err))
}
- ctx, cancel := context.WithTimeout(context.Background(), 5*time.Minute)
- defer cancel()
+ cfg := reindexConfig{Batch: reindexBatch, Timeout: reindexTimeout, Full: reindexFull}
+ cached, err := cachedEmbeddings(v, emb.Name(), cfg)
+ if err != nil {
+ return errors.Join(ErrVault, err)
+ }
+
idx := &index.Index{Provider: emb.Name(), BuiltAt: time.Now()}
- for start := 0; start < len(entries); start += reindexBatch {
- end := min(start+reindexBatch, len(entries))
- batch := entries[start:end]
- texts := make([]string, len(batch))
- for i, e := range batch {
- texts[i] = e.Body
- }
- vecs, err := emb.Embed(ctx, texts)
- if err != nil {
- return errors.Join(ErrLLM, fmt.Errorf("embed batch: %w", err))
- }
- for i, vec := range vecs {
- idx.Entries = append(idx.Entries, index.Entry{
- Time: batch[i].Time,
- Tags: batch[i].Tags,
- Body: batch[i].Body,
- Embedding: vec,
- })
+ idx.Entries = make([]index.Entry, 0, len(entries))
+ pending := make([]int, 0, len(entries))
+ for _, e := range entries {
+ entry := index.Entry{Time: e.Time, Tags: e.Tags, Body: e.Body}
+ if vec, ok := cached[index.ContentHash(e.Body)]; ok {
+ entry.Embedding = vec
+ } else {
+ pending = append(pending, len(idx.Entries))
}
- fmt.Fprintf(cmd.ErrOrStderr(), "Embedded %d/%d\n", end, len(entries))
+ idx.Entries = append(idx.Entries, entry)
}
- if len(idx.Entries) > 0 {
- idx.Dim = len(idx.Entries[0].Embedding)
+ reused := len(idx.Entries) - len(pending)
+ fmt.Fprintf(cmd.ErrOrStderr(), "Indexing %d entries: %d reused, %d to embed via %s\n",
+ len(idx.Entries), reused, len(pending), emb.Name())
+
+ if err := embedPending(cmd, v, emb, idx, pending, cfg); err != nil {
+ return err
}
+ setDim(idx)
if err := saveIndex(v, idx); err != nil {
return errors.Join(ErrVault, err)
}
- fmt.Fprintf(cmd.OutOrStdout(), "Wrote index with %d entries via %s\n", len(idx.Entries), emb.Name())
+ fmt.Fprintf(cmd.OutOrStdout(), "Wrote index with %d entries via %s (%d reused, %d embedded)\n",
+ len(idx.Entries), emb.Name(), reused, len(pending))
+ return nil
+}
+
+// embedPending embeds the entries at the given positions in batches, writing a
+// checkpoint as it goes. A failed batch leaves the checkpoint in place so the
+// embeddings already bought are not lost.
+func embedPending(
+ cmd *cobra.Command,
+ v *vault.Vault,
+ emb llm.Embedder,
+ idx *index.Index,
+ pending []int,
+ cfg reindexConfig,
+) error {
+ done := 0
+ for start := 0; start < len(pending); start += cfg.Batch {
+ end := min(start+cfg.Batch, len(pending))
+ positions := pending[start:end]
+ texts := make([]string, len(positions))
+ for i, pos := range positions {
+ texts[i] = idx.Entries[pos].Body
+ }
+ vecs, err := embedBatch(emb, texts, cfg.Timeout)
+ if err != nil {
+ checkpoint(cmd, v, idx)
+ return errors.Join(ErrLLM, fmt.Errorf("embed entries %d-%d of %d: %w", start+1, end, len(pending), err))
+ }
+ if len(vecs) != len(positions) {
+ checkpoint(cmd, v, idx)
+ return errors.Join(ErrLLM, fmt.Errorf("embedder returned %d vectors for %d entries", len(vecs), len(positions)))
+ }
+ for i, pos := range positions {
+ idx.Entries[pos].Embedding = vecs[i]
+ }
+ done = end
+ fmt.Fprintf(cmd.ErrOrStderr(), "Embedded %d/%d\n", done, len(pending))
+ if batchNum := (start / cfg.Batch) + 1; batchNum%reindexCheckpointEvery == 0 {
+ checkpoint(cmd, v, idx)
+ }
+ }
return nil
}
+// embedBatch runs one provider call under its own deadline.
+func embedBatch(emb llm.Embedder, texts []string, timeout time.Duration) ([][]float32, error) {
+ ctx, cancel := context.WithTimeout(context.Background(), timeout)
+ defer cancel()
+ return emb.Embed(ctx, texts)
+}
+
+// checkpoint writes the index as it currently stands, dropping entries with no
+// vector yet so what lands on disk is a smaller but valid index. A failure to
+// write is reported and otherwise ignored, since the caller is either mid-run
+// or already returning a more useful error.
+func checkpoint(cmd *cobra.Command, v *vault.Vault, idx *index.Index) {
+ partial := &index.Index{Provider: idx.Provider, BuiltAt: idx.BuiltAt}
+ for _, e := range idx.Entries {
+ if len(e.Embedding) > 0 {
+ partial.Entries = append(partial.Entries, e)
+ }
+ }
+ if len(partial.Entries) == 0 {
+ return
+ }
+ setDim(partial)
+ if err := saveIndex(v, partial); err != nil {
+ fmt.Fprintf(cmd.ErrOrStderr(), "Warning: could not checkpoint the index: %v\n", err)
+ return
+ }
+ fmt.Fprintf(cmd.ErrOrStderr(), "Checkpointed %d embedded entries.\n", len(partial.Entries))
+}
+
+// cachedEmbeddings returns the vectors already on disk for reuse, or nothing
+// when a full rebuild was asked for or the stored index came from a different
+// embedder. Vectors from another provider have their own geometry and
+// dimension, so mixing them would silently corrupt every later search.
+func cachedEmbeddings(v *vault.Vault, provider string, cfg reindexConfig) (map[string][]float32, error) {
+ if cfg.Full {
+ return nil, nil
+ }
+ idx, err := loadIndex(v)
+ if err != nil {
+ return nil, fmt.Errorf("load index: %w", err)
+ }
+ if idx.Provider != provider {
+ return nil, nil
+ }
+ return idx.EmbeddingsByContent(), nil
+}
+
+// setDim records the embedding width from the first vector present.
+func setDim(idx *index.Index) {
+ for _, e := range idx.Entries {
+ if len(e.Embedding) > 0 {
+ idx.Dim = len(e.Embedding)
+ return
+ }
+ }
+}
+
// allEntries returns every entry across the vault in chronological order.
func allEntries(v *vault.Vault) ([]vault.Entry, error) {
days, err := v.ListDays()
diff --git a/cmd/cmd_reindex_test.go b/cmd/cmd_reindex_test.go
new file mode 100644
index 0000000..47cd556
--- /dev/null
+++ b/cmd/cmd_reindex_test.go
@@ -0,0 +1,203 @@
+package cmd
+
+import (
+ "context"
+ "errors"
+ "strings"
+ "sync/atomic"
+ "testing"
+ "time"
+
+ "github.com/spf13/cobra"
+
+ "github.com/dcadolph/midden/index"
+ "github.com/dcadolph/midden/internal/vault"
+)
+
+// mockEmbedder is an Embedder whose vectors are produced by a configurable
+// function, recording how many texts it was asked to embed.
+type mockEmbedder struct {
+ // EmbedFunc produces the vectors for one call.
+ EmbedFunc func(texts []string) ([][]float32, error)
+ // embedded counts the texts passed across every call.
+ embedded atomic.Int64
+ // calls counts how many times Embed was invoked.
+ calls atomic.Int64
+}
+
+// Embed delegates to EmbedFunc, recording the volume it was asked for.
+func (m *mockEmbedder) Embed(_ context.Context, texts []string) ([][]float32, error) {
+ m.calls.Add(1)
+ m.embedded.Add(int64(len(texts)))
+ return m.EmbedFunc(texts)
+}
+
+// Dim reports the fixed width of the mock's vectors.
+func (m *mockEmbedder) Dim() int { return 2 }
+
+// Name identifies the mock provider.
+func (m *mockEmbedder) Name() string { return "mock:embed" }
+
+// unitVectors returns one distinct two-dimensional vector per input.
+func unitVectors(texts []string) ([][]float32, error) {
+ out := make([][]float32, len(texts))
+ for i, t := range texts {
+ out[i] = []float32{float32(len(t)), 1}
+ }
+ return out, nil
+}
+
+// testReindexConfig returns a rebuild config with a generous per-call deadline,
+// so a slow machine never turns a unit test into a timeout failure.
+func testReindexConfig(batch int) reindexConfig {
+ return reindexConfig{Batch: batch, Timeout: time.Minute}
+}
+
+// reindexVault returns a temp vault and a command with output discarded.
+func reindexVault(t *testing.T) (*vault.Vault, *cobra.Command) {
+ t.Helper()
+ v, err := vault.Open(t.TempDir())
+ if err != nil {
+ t.Fatalf("vault.Open: %v", err)
+ }
+ c := &cobra.Command{}
+ c.SetOut(new(strings.Builder))
+ c.SetErr(new(strings.Builder))
+ return v, c
+}
+
+// buildIndex assembles an index over the given bodies with no embeddings yet,
+// plus the positions needing one.
+func buildIndex(bodies ...string) (*index.Index, []int) {
+ idx := &index.Index{Provider: "mock:embed"}
+ pending := make([]int, 0, len(bodies))
+ for i, b := range bodies {
+ idx.Entries = append(idx.Entries, index.Entry{Body: b})
+ pending = append(pending, i)
+ }
+ return idx, pending
+}
+
+func TestEmbedPendingFillsEveryPendingEntry(t *testing.T) {
+ t.Parallel()
+ v, cmd := reindexVault(t)
+ idx, pending := buildIndex("alpha", "beta", "gamma", "delta", "epsilon")
+ emb := &mockEmbedder{EmbedFunc: unitVectors}
+ if err := embedPending(cmd, v, emb, idx, pending, testReindexConfig(2)); err != nil {
+ t.Fatalf("embedPending: %v", err)
+ }
+ for i, e := range idx.Entries {
+ if len(e.Embedding) == 0 {
+ t.Errorf("entry %d (%q) was left without an embedding", i, e.Body)
+ }
+ }
+ if got := emb.embedded.Load(); got != 5 {
+ t.Errorf("want 5 texts embedded, got %d", got)
+ }
+ // Five entries at a batch size of two is three calls, which is what keeps a
+ // large backfill inside a per-call deadline instead of one run-long one.
+ if got := emb.calls.Load(); got != 3 {
+ t.Errorf("want 3 batched calls, got %d", got)
+ }
+}
+
+func TestEmbedPendingCheckpointsBeforeReturningAFailure(t *testing.T) {
+ t.Parallel()
+ v, cmd := reindexVault(t)
+ idx, pending := buildIndex("alpha", "beta", "gamma", "delta")
+ emb := &mockEmbedder{EmbedFunc: func(texts []string) ([][]float32, error) {
+ if strings.Contains(strings.Join(texts, ","), "gamma") {
+ return nil, errors.New("provider exploded")
+ }
+ return unitVectors(texts)
+ }}
+ err := embedPending(cmd, v, emb, idx, pending, testReindexConfig(2))
+ if err == nil {
+ t.Fatal("want an error when a batch fails")
+ }
+ // The first batch was paid for, so it has to survive the failure or the next
+ // run buys the same embeddings again.
+ saved, err := loadIndex(v)
+ if err != nil {
+ t.Fatalf("loadIndex: %v", err)
+ }
+ if len(saved.Entries) != 2 {
+ t.Errorf("want the 2 completed entries checkpointed, got %d", len(saved.Entries))
+ }
+ if saved.Dim != 2 {
+ t.Errorf("want the checkpoint to record the embedding width, got %d", saved.Dim)
+ }
+}
+
+func TestEmbedPendingRejectsAShortProviderResponse(t *testing.T) {
+ t.Parallel()
+ v, cmd := reindexVault(t)
+ idx, pending := buildIndex("alpha", "beta", "gamma")
+ emb := &mockEmbedder{EmbedFunc: func(texts []string) ([][]float32, error) {
+ vecs, _ := unitVectors(texts)
+ return vecs[:len(vecs)-1], nil
+ }}
+ // Silently accepting fewer vectors than texts would shift every embedding
+ // onto the wrong entry.
+ if err := embedPending(cmd, v, emb, idx, pending, testReindexConfig(4)); err == nil {
+ t.Fatal("want an error when the provider returns fewer vectors than texts")
+ }
+}
+
+func TestCachedEmbeddingsReuseAndInvalidation(t *testing.T) {
+ t.Parallel()
+ v, _ := reindexVault(t)
+ stored := &index.Index{
+ Provider: "mock:embed",
+ Dim: 2,
+ Entries: []index.Entry{
+ {Body: "alpha", Embedding: []float32{1, 0}},
+ {Body: "beta", Embedding: []float32{0, 1}},
+ },
+ }
+ if err := saveIndex(v, stored); err != nil {
+ t.Fatalf("saveIndex: %v", err)
+ }
+
+ cached, err := cachedEmbeddings(v, "mock:embed", testReindexConfig(4))
+ if err != nil {
+ t.Fatalf("cachedEmbeddings: %v", err)
+ }
+ if len(cached) != 2 {
+ t.Errorf("want 2 reusable vectors, got %d", len(cached))
+ }
+ if _, ok := cached[index.ContentHash("alpha")]; !ok {
+ t.Error("want the vector for unchanged text to be reusable")
+ }
+
+ // A different provider means a different geometry, so nothing may carry over.
+ other, err := cachedEmbeddings(v, "mock:other", testReindexConfig(4))
+ if err != nil {
+ t.Fatalf("cachedEmbeddings: %v", err)
+ }
+ if len(other) != 0 {
+ t.Errorf("want no reuse across providers, got %d vectors", len(other))
+ }
+
+ fullCfg := testReindexConfig(4)
+ fullCfg.Full = true
+ full, err := cachedEmbeddings(v, "mock:embed", fullCfg)
+ if err != nil {
+ t.Fatalf("cachedEmbeddings: %v", err)
+ }
+ if len(full) != 0 {
+ t.Errorf("want --full to ignore the cache, got %d vectors", len(full))
+ }
+}
+
+func TestSetDimSkipsEntriesWithoutVectors(t *testing.T) {
+ t.Parallel()
+ idx := &index.Index{Entries: []index.Entry{
+ {Body: "no vector yet"},
+ {Body: "embedded", Embedding: []float32{1, 2, 3}},
+ }}
+ setDim(idx)
+ if idx.Dim != 3 {
+ t.Errorf("want dim 3, got %d", idx.Dim)
+ }
+}
diff --git a/cmd/cmd_stats.go b/cmd/cmd_stats.go
index 2c37d5b..d54fc92 100644
--- a/cmd/cmd_stats.go
+++ b/cmd/cmd_stats.go
@@ -1,6 +1,8 @@
package cmd
import (
+ "time"
+
"errors"
"fmt"
@@ -30,7 +32,7 @@ func runStats(cmd *cobra.Command, _ []string) error {
if err != nil {
return err
}
- s, err := v.ComputeStats(statsTopTags)
+ s, err := v.ComputeStats(statsTopTags, time.Now())
if err != nil {
return errors.Join(ErrVault, fmt.Errorf("compute stats: %w", err))
}
@@ -54,7 +56,20 @@ func runStats(cmd *cobra.Command, _ []string) error {
if !s.FirstEntry.IsZero() {
heading("Span")
fmt.Fprintf(w, " First: %s\n", s.FirstEntry.Format(layoutDateTime))
- fmt.Fprintf(w, " Last: %s\n", s.LastEntry.Format(layoutDateTime))
+ last := s.LastPast
+ if last.IsZero() {
+ last = s.LastEntry
+ }
+ fmt.Fprintf(w, " Last: %s\n", last.Format(layoutDateTime))
+ if s.Scheduled > 0 {
+ fmt.Fprintf(w, " Ahead: %d scheduled, through %s\n", s.Scheduled, s.LastEntry.Format(layoutDate))
+ }
+ }
+ heading("Yours")
+ if s.Authored == 0 {
+ fmt.Fprintln(w, " Written by you: 0 entries. Everything here so far was imported.")
+ } else {
+ fmt.Fprintf(w, " Written by you: %d entries, last on %s\n", s.Authored, s.LastAuthored.Format(layoutDate))
}
if len(s.TopTags) > 0 {
heading("Top tags")
diff --git a/cmd/cmd_streak.go b/cmd/cmd_streak.go
index 8a138a0..8f052eb 100644
--- a/cmd/cmd_streak.go
+++ b/cmd/cmd_streak.go
@@ -5,6 +5,7 @@ import (
"fmt"
"time"
+ "github.com/dcadolph/midden/internal/vault"
"github.com/spf13/cobra"
)
@@ -25,7 +26,9 @@ func runStreak(cmd *cobra.Command, _ []string) error {
if err != nil {
return err
}
- n, err := v.Streak(time.Now())
+ // Only days the person actually wrote something count. Imported calendar
+ // events would otherwise report a streak for appointments merely attended.
+ n, err := v.Streak(time.Now(), vault.Entry.Authored)
if err != nil {
return errors.Join(ErrVault, fmt.Errorf("streak: %w", err))
}
diff --git a/cmd/cmd_undo.go b/cmd/cmd_undo.go
index 2ac610c..0e91f8a 100644
--- a/cmd/cmd_undo.go
+++ b/cmd/cmd_undo.go
@@ -1,6 +1,8 @@
package cmd
import (
+ "time"
+
"bytes"
"errors"
"fmt"
@@ -36,7 +38,15 @@ func runUndo(cmd *cobra.Command, _ []string) error {
if err != nil {
return errors.Join(ErrVault, fmt.Errorf("list days: %w", err))
}
+ // Undo removes the last thing the person wrote, and nothing they wrote lives
+ // in the future. An imported calendar does: without this cutoff the walk
+ // starts at next year's appointments and deletes one of those instead, which
+ // is the single most destructive way the scheduled-entry problem can land.
+ today := dayStart(time.Now())
for i := len(days) - 1; i >= 0; i-- {
+ if days[i].After(today) {
+ continue
+ }
path := v.DayPath(days[i])
data, err := v.ReadBytes(path)
if err != nil {
diff --git a/cmd/cmd_weave.go b/cmd/cmd_weave.go
new file mode 100644
index 0000000..0b16a13
--- /dev/null
+++ b/cmd/cmd_weave.go
@@ -0,0 +1,320 @@
+package cmd
+
+import (
+ "errors"
+ "fmt"
+ "io"
+ "strings"
+ "time"
+
+ "github.com/spf13/cobra"
+
+ "github.com/dcadolph/midden/internal/jsonutil"
+ "github.com/dcadolph/midden/internal/vault"
+ "github.com/dcadolph/midden/internal/weave"
+)
+
+// Weave options.
+var (
+ weaveSince string
+ weaveUntil string
+ weaveMin int
+ weaveLimit int
+ weaveSources []string
+ weaveTags []string
+ weaveExplain bool
+)
+
+// weaveCmd surfaces the shape of the record over time.
+var weaveCmd = &cobra.Command{
+ Use: "weave",
+ Short: "Show what recurs in the record, when it started, and when it stopped.",
+ Long: "Weave finds the threads running through the vault.\n\n" +
+ "Search answers what you already know to ask about. Weave answers what you cannot ask, " +
+ "because a person can recall what they did but cannot perceive absence: nothing marks the " +
+ "last time something happened. Every figure here is counted rather than inferred, so there " +
+ "is no model in the path and nothing to invent.",
+ RunE: runWeave,
+}
+
+func init() {
+ weaveCmd.Flags().StringVar(&weaveSince, "since", "", "Only consider entries on or after this date.")
+ weaveCmd.Flags().StringVar(&weaveUntil, "until", "", "Only consider entries on or before this date.")
+ weaveCmd.Flags().IntVar(&weaveMin, "min", 5, "Fewest occurrences a pattern needs to count as a thread.")
+ weaveCmd.Flags().IntVar(&weaveLimit, "top", 12, "Maximum threads to show per section.")
+ weaveCmd.Flags().BoolVar(&weaveExplain, "explain", false,
+ "Show the distinct headlines folded into each thread, so a grouping can be checked before it is believed.")
+ weaveCmd.Flags().StringSliceVar(&weaveTags, "tag", nil,
+ "Only weave entries carrying one of these tags. Commit history repeats boilerplate subjects across repositories, so restricting to a life source such as calendar keeps those out of the threads.")
+ weaveCmd.Flags().StringSliceVar(&weaveSources, "source", []string{"calendar", "git"},
+ "Source tags to compare when looking for days where parts of your life meet.")
+ rootCmd.AddCommand(weaveCmd)
+}
+
+// runWeave reads the vault, detects threads, and reports what ended, what began,
+// and where separate parts of the record meet.
+func runWeave(cmd *cobra.Command, _ []string) error {
+ span, err := resolveDateRange(weaveSince, weaveUntil)
+ if err != nil {
+ return err
+ }
+ v, err := openVault()
+ if err != nil {
+ return err
+ }
+ entries, err := entriesInRange(v, span)
+ if err != nil {
+ return errors.Join(ErrVault, fmt.Errorf("scan vault: %w", err))
+ }
+ if len(entries) == 0 {
+ return errors.Join(ErrNotFound, fmt.Errorf("no entries in range (%s)", span.Label()))
+ }
+
+ now := time.Now()
+ threadInput := entries
+ if len(weaveTags) > 0 {
+ threadInput = filterByTag(entries, weaveTags)
+ if len(threadInput) == 0 {
+ return errors.Join(ErrNotFound, fmt.Errorf("no entries carrying %s", strings.Join(weaveTags, ", ")))
+ }
+ }
+ opts := weave.DefaultOptions(now)
+ opts.MinCount = weaveMin
+ threads := weave.Threads(threadInput, opts)
+ handoffs := weave.Handoffs(threads, weave.DefaultHandoffOptions())
+ overlaps := weave.Overlaps(entries, weaveSources, now)
+ gaps := weave.GapsBySource(entries, weaveSources, weave.DefaultGapOptions(now))
+
+ if jsonOutput {
+ return jsonutil.Encode(cmd.OutOrStdout(), weaveJSON{
+ Threads: threadsToJSON(threads),
+ Handoffs: handoffsToJSON(handoffs),
+ }, jsonPretty)
+ }
+ writeWeave(cmd.OutOrStdout(), threads, handoffs, overlaps, gaps, weaveLimit)
+ return nil
+}
+
+// writeWeave renders the human-readable report.
+func writeWeave(
+ w io.Writer,
+ threads []weave.Thread,
+ handoffs []weave.Handoff,
+ overlaps []weave.Overlap,
+ gaps []weave.Gap,
+ limit int,
+) {
+ byStatus := func(s weave.Status) []weave.Thread {
+ var out []weave.Thread
+ for _, t := range threads {
+ if t.Status == s {
+ out = append(out, t)
+ }
+ }
+ return out
+ }
+
+ if len(gaps) > 0 {
+ section(w, "Silences", "stretches where the record itself went quiet")
+ for _, g := range head(gaps, limit) {
+ scope := "whole record"
+ if g.Source != "" {
+ scope = g.Source + " only"
+ }
+ fmt.Fprintf(w, " %s to %s %d months, %d entries [%s] (about %.0f/month before, %.0f after)\n",
+ g.From.Format("2006-01"), g.To.Format("2006-01"), g.Months, g.Entries, scope, g.Before, g.After)
+ }
+ }
+
+ ended := byStatus(weave.Ended)
+ section(w, "Ended", "things that stopped without anything marking the last one")
+ for _, t := range head(ended, limit) {
+ fmt.Fprintf(w, " %-44s %4dx over %4s last %s, %s ago\n",
+ truncate(t.Label, 44), t.Count, years(t.SpanDays), t.Last.Format(layoutDate), months(t.SilentDays))
+ }
+ if len(ended) == 0 {
+ fmt.Fprintln(w, " nothing has gone quiet")
+ }
+
+ dormant := byStatus(weave.Dormant)
+ if len(dormant) > 0 {
+ section(w, "Between seasons", "quiet, but they have come back from a gap this long before")
+ for _, t := range head(dormant, limit) {
+ fmt.Fprintf(w, " %-44s %4dx, last %s, quiet %s (longest gap before: %s)\n",
+ truncate(t.Label, 44), t.Count, t.Last.Format(layoutDate),
+ months(t.SilentDays), months(t.MaxGap))
+ writeVariants(w, t)
+ }
+ }
+
+ emerging := byStatus(weave.Emerging)
+ section(w, "Started", "threads that began recently")
+ for _, t := range head(emerging, limit) {
+ fmt.Fprintf(w, " %-44s %4dx since %s\n", truncate(t.Label, 44), t.Count, t.First.Format(layoutDate))
+ }
+ if len(emerging) == 0 {
+ fmt.Fprintln(w, " nothing new")
+ }
+
+ ongoing := byStatus(weave.Ongoing)
+ section(w, "Ongoing", "the steady weight of the record")
+ for _, t := range head(ongoing, limit) {
+ fmt.Fprintf(w, " %-44s %4dx every ~%d days\n", truncate(t.Label, 44), t.Count, t.MedianGap)
+ writeVariants(w, t)
+ }
+ if len(ongoing) == 0 {
+ fmt.Fprintln(w, " nothing recurring")
+ }
+
+ if len(handoffs) > 0 {
+ section(w, "Handoffs", "one thread ended and another began soon after")
+ for _, h := range head(handoffs, limit) {
+ fmt.Fprintf(w, " %s\n", truncate(h.From.Label, 60))
+ fmt.Fprintf(w, " ended %s after %dx, then %s began %d days later\n",
+ h.From.Last.Format(layoutDate), h.From.Count, truncate(h.To.Label, 40), h.GapDays)
+ }
+ }
+
+ if len(overlaps) > 0 {
+ section(w, "Crossings", "days where separate parts of the record meet")
+ for _, o := range head(overlaps, limit) {
+ parts := make([]string, 0, len(o.Counts))
+ for src, n := range o.Counts {
+ parts = append(parts, fmt.Sprintf("%d %s", n, src))
+ }
+ fmt.Fprintf(w, " %s (%s)\n", o.Day.Format(layoutDate), strings.Join(parts, ", "))
+ for _, src := range weaveSources {
+ if h := o.Headlines[src]; h != "" {
+ fmt.Fprintf(w, " %-9s %s\n", src+":", truncate(h, 62))
+ }
+ }
+ }
+ }
+}
+
+// writeVariants lists the headlines folded into a thread when explaining. A
+// claim that something ended rests entirely on what was grouped together, so
+// the grouping has to be inspectable.
+func writeVariants(w io.Writer, t weave.Thread) {
+ if !weaveExplain || len(t.Variants) < 2 {
+ return
+ }
+ for _, v := range t.Variants {
+ fmt.Fprintf(w, " · %s\n", truncate(v, 68))
+ }
+}
+
+// section writes a titled block header.
+func section(w io.Writer, title, blurb string) {
+ fmt.Fprintf(w, "\n%s — %s\n\n", title, blurb)
+}
+
+// head returns at most n items from the slice.
+func head[T any](items []T, n int) []T {
+ if n > 0 && len(items) > n {
+ return items[:n]
+ }
+ return items
+}
+
+// years renders a day count as an approximate number of years.
+func years(days int) string {
+ return fmt.Sprintf("%.1fy", float64(days)/365)
+}
+
+// months renders a day count as an approximate number of months.
+func months(days int) string {
+ if days < 60 {
+ return fmt.Sprintf("%d days", days)
+ }
+ return fmt.Sprintf("%d months", days/30)
+}
+
+// truncate shortens a label to fit a column.
+func truncate(s string, n int) string {
+ if len(s) <= n {
+ return s
+ }
+ return s[:n-1] + "…"
+}
+
+// filterByTag keeps only entries carrying one of the given tags.
+func filterByTag(entries []vault.Entry, tags []string) []vault.Entry {
+ want := make(map[string]bool, len(tags))
+ for _, t := range tags {
+ want[strings.ToLower(strings.TrimPrefix(t, "#"))] = true
+ }
+ var out []vault.Entry
+ for _, e := range entries {
+ for _, t := range e.Tags {
+ if want[strings.ToLower(t)] {
+ out = append(out, e)
+ break
+ }
+ }
+ }
+ return out
+}
+
+// entriesInRange reads every entry inside the window.
+func entriesInRange(v *vault.Vault, span dateRange) ([]vault.Entry, error) {
+ var out []vault.Entry
+ err := forEachEntryInRange(v, span, func(e vault.Entry) { out = append(out, e) })
+ return out, err
+}
+
+// weaveJSON is the wire shape for structured weave output.
+type weaveJSON struct {
+ // Threads are the recurring patterns found in the record.
+ Threads []threadJSON `json:"threads"`
+ // Handoffs are the successions between threads.
+ Handoffs []handoffJSON `json:"handoffs,omitempty"`
+}
+
+// threadJSON is the wire shape of one thread.
+type threadJSON struct {
+ // Label is the representative title.
+ Label string `json:"label"`
+ // Status is where the thread stands.
+ Status string `json:"status"`
+ // Count is how many occurrences it holds.
+ Count int `json:"count"`
+ // First and Last are the bounding dates.
+ First string `json:"first"`
+ Last string `json:"last"`
+ // MedianGapDays is the typical spacing between occurrences.
+ MedianGapDays int `json:"median_gap_days"`
+ // SilentDays is how long it has been quiet.
+ SilentDays int `json:"silent_days"`
+}
+
+// handoffJSON is the wire shape of one succession.
+type handoffJSON struct {
+ // From and To are the labels of the ended and begun threads.
+ From string `json:"from"`
+ To string `json:"to"`
+ // GapDays is how long passed between them.
+ GapDays int `json:"gap_days"`
+}
+
+// threadsToJSON converts threads to their wire shape.
+func threadsToJSON(threads []weave.Thread) []threadJSON {
+ out := make([]threadJSON, len(threads))
+ for i, t := range threads {
+ out[i] = threadJSON{
+ Label: t.Label, Status: string(t.Status), Count: t.Count,
+ First: t.First.Format(layoutDate), Last: t.Last.Format(layoutDate),
+ MedianGapDays: t.MedianGap, SilentDays: t.SilentDays,
+ }
+ }
+ return out
+}
+
+// handoffsToJSON converts successions to their wire shape.
+func handoffsToJSON(handoffs []weave.Handoff) []handoffJSON {
+ out := make([]handoffJSON, len(handoffs))
+ for i, h := range handoffs {
+ out[i] = handoffJSON{From: h.From.Label, To: h.To.Label, GapDays: h.GapDays}
+ }
+ return out
+}
diff --git a/cmd/daterange.go b/cmd/daterange.go
new file mode 100644
index 0000000..96fc778
--- /dev/null
+++ b/cmd/daterange.go
@@ -0,0 +1,60 @@
+package cmd
+
+import (
+ "fmt"
+ "time"
+
+ "github.com/dcadolph/midden/dateutil"
+)
+
+// dateRange bounds a query to a span of days. A zero bound is open.
+type dateRange struct {
+ // From is the inclusive start of the range at local midnight.
+ From time.Time
+ // To is the inclusive end of the range at the last instant of that day.
+ To time.Time
+}
+
+// Bounded reports whether either end of the range is set.
+func (r dateRange) Bounded() bool {
+ return !r.From.IsZero() || !r.To.IsZero()
+}
+
+// Label renders the range for human-readable output.
+func (r dateRange) Label() string {
+ switch {
+ case !r.Bounded():
+ return "whole vault"
+ case r.From.IsZero():
+ return "through " + r.To.Format(layoutDate)
+ case r.To.IsZero():
+ return "since " + r.From.Format(layoutDate)
+ }
+ return r.From.Format(layoutDate) + " to " + r.To.Format(layoutDate)
+}
+
+// resolveDateRange parses the since and until flag values. Both accept every
+// form dateutil understands, so "2026-08-01", "30-days-ago", and "monday" are
+// equivalent kinds of input. The until bound extends to the end of its day so a
+// single-day range still covers that day's entries.
+func resolveDateRange(since, until string) (dateRange, error) {
+ var r dateRange
+ if since != "" {
+ from, err := dateutil.Parse(since)
+ if err != nil {
+ return dateRange{}, fmt.Errorf("parse --since %q: %w", since, err)
+ }
+ r.From = from
+ }
+ if until != "" {
+ to, err := dateutil.Parse(until)
+ if err != nil {
+ return dateRange{}, fmt.Errorf("parse --until %q: %w", until, err)
+ }
+ r.To = to.AddDate(0, 0, 1).Add(-time.Nanosecond)
+ }
+ if !r.From.IsZero() && !r.To.IsZero() && r.To.Before(r.From) {
+ return dateRange{}, fmt.Errorf("--until %q is before --since %q", until, since)
+ }
+ return r, nil
+}
diff --git a/cmd/daterange_test.go b/cmd/daterange_test.go
new file mode 100644
index 0000000..4daff1c
--- /dev/null
+++ b/cmd/daterange_test.go
@@ -0,0 +1,110 @@
+package cmd
+
+import (
+ "fmt"
+ "testing"
+ "time"
+
+ "github.com/google/go-cmp/cmp"
+)
+
+func TestResolveDateRange(t *testing.T) {
+ t.Parallel()
+ tests := []struct {
+ Since string
+ Until string
+ WantLabel string
+ WantFrom string
+ WantTo string
+ Want bool
+ WantBound bool
+ }{{ // Test 0: Both ends empty leaves the range open.
+ WantLabel: "whole vault",
+ WantBound: false,
+ }, { // Test 1: A since bound alone leaves the far end open.
+ Since: "2024-03-01",
+ WantFrom: "2024-03-01 00:00:00",
+ WantLabel: "since 2024-03-01",
+ WantBound: true,
+ }, { // Test 2: An until bound extends to the last instant of its day.
+ Until: "2024-03-31",
+ WantTo: "2024-03-31 23:59:59",
+ WantLabel: "through 2024-03-31",
+ WantBound: true,
+ }, { // Test 3: Both bounds render as a span.
+ Since: "2024-03-01",
+ Until: "2024-03-31",
+ WantFrom: "2024-03-01 00:00:00",
+ WantTo: "2024-03-31 23:59:59",
+ WantLabel: "2024-03-01 to 2024-03-31",
+ WantBound: true,
+ }, { // Test 4: A single day covers that whole day.
+ Since: "2024-03-15",
+ Until: "2024-03-15",
+ WantFrom: "2024-03-15 00:00:00",
+ WantTo: "2024-03-15 23:59:59",
+ WantLabel: "2024-03-15 to 2024-03-15",
+ WantBound: true,
+ }, { // Test 5: An inverted range is rejected rather than silently returning nothing.
+ Since: "2024-03-31",
+ Until: "2024-03-01",
+ Want: true,
+ }, { // Test 6: An unparseable since value is rejected.
+ Since: "not-a-date",
+ Want: true,
+ }, { // Test 7: An unparseable until value is rejected.
+ Until: "someday",
+ Want: true,
+ }}
+ for testNum, test := range tests {
+ t.Run(fmt.Sprintf("test %d", testNum), func(t *testing.T) {
+ t.Parallel()
+ got, err := resolveDateRange(test.Since, test.Until)
+ if test.Want {
+ if err == nil {
+ t.Fatalf("want error, got range %+v", got)
+ }
+ return
+ }
+ if err != nil {
+ t.Fatalf("resolveDateRange: %v", err)
+ }
+ if diff := cmp.Diff(test.WantBound, got.Bounded()); diff != "" {
+ t.Errorf("Bounded mismatch (-want +got):\n%s", diff)
+ }
+ if diff := cmp.Diff(test.WantLabel, got.Label()); diff != "" {
+ t.Errorf("Label mismatch (-want +got):\n%s", diff)
+ }
+ if diff := cmp.Diff(test.WantFrom, formatBound(got.From)); diff != "" {
+ t.Errorf("From mismatch (-want +got):\n%s", diff)
+ }
+ if diff := cmp.Diff(test.WantTo, formatBound(got.To)); diff != "" {
+ t.Errorf("To mismatch (-want +got):\n%s", diff)
+ }
+ })
+ }
+}
+
+func TestResolveDateRangeUntilCoversTrailingSecond(t *testing.T) {
+ t.Parallel()
+ got, err := resolveDateRange("", "2024-03-31")
+ if err != nil {
+ t.Fatalf("resolveDateRange: %v", err)
+ }
+ last := time.Date(2024, time.March, 31, 23, 59, 59, 999999999, time.Local)
+ if got.To.Before(last) {
+ t.Errorf("until bound %s excludes the final instant %s of its day", got.To, last)
+ }
+ if !got.To.Before(time.Date(2024, time.April, 1, 0, 0, 0, 0, time.Local)) {
+ t.Errorf("until bound %s spills into the next day", got.To)
+ }
+}
+
+// formatBound renders a range bound for comparison, using an empty string for
+// the zero time so an open end is distinguishable from a real timestamp.
+func formatBound(t time.Time) string {
+ if t.IsZero() {
+ return ""
+ }
+ return t.Format(layoutDateTime)
+}
diff --git a/ics/expand.go b/ics/expand.go
new file mode 100644
index 0000000..0876a0e
--- /dev/null
+++ b/ics/expand.go
@@ -0,0 +1,129 @@
+package ics
+
+import (
+ "sort"
+ "time"
+)
+
+// ExpandReport records what expansion could not do faithfully, so a caller can
+// say so rather than presenting a partial calendar as a complete one.
+type ExpandReport struct {
+ // Occurrences is the number of events produced from recurring series.
+ Occurrences int
+ // Unexpanded is the number of recurring events whose rule midden does not
+ // expand, each of which contributes only its first occurrence.
+ Unexpanded int
+ // Truncated is the number of series whose expansion hit the period cap and
+ // may therefore be missing later occurrences.
+ Truncated int
+ // Excluded is the number of generated occurrences dropped by EXDATE.
+ Excluded int
+ // Overridden is the number of generated occurrences replaced by an explicit
+ // override event carrying the same recurrence identifier.
+ Overridden int
+}
+
+// Expand turns parsed events into the concrete occurrences falling inside the
+// window, ordered by start time. A recurring event becomes one event per
+// occurrence with its start and end shifted and its UID preserved, so a series
+// contributes every time it actually happened rather than only the first. A
+// zero from or to leaves that end of the window open, though a rule bounded by
+// neither the window nor its own COUNT or UNTIL is reported as truncated rather
+// than walked forever.
+func Expand(events []Event, from, to time.Time) ([]Event, ExpandReport) {
+ var report ExpandReport
+ overrides := overrideIndex(events)
+ var out []Event
+ for _, e := range events {
+ // An override is already a standalone event; it is emitted on its own
+ // terms and suppresses the occurrence it replaces.
+ if !e.RecurrenceID.IsZero() || !e.Recurs() {
+ if within(e.Start, from, to) {
+ out = append(out, e)
+ }
+ continue
+ }
+ if e.Rule == nil {
+ report.Unexpanded++
+ if within(e.Start, from, to) {
+ out = append(out, e)
+ }
+ continue
+ }
+ starts, truncated := e.Rule.Occurrences(e.Start, from, to)
+ if truncated {
+ report.Truncated++
+ }
+ for _, start := range starts {
+ if excluded(e.ExDates, start) {
+ report.Excluded++
+ continue
+ }
+ if overrides[e.UID][start.UnixNano()] {
+ report.Overridden++
+ continue
+ }
+ report.Occurrences++
+ out = append(out, occurrence(e, start))
+ }
+ }
+ sort.SliceStable(out, func(i, j int) bool { return out[i].Start.Before(out[j].Start) })
+ return out, report
+}
+
+// occurrence returns a copy of the series event moved to the given start,
+// holding its duration. The rule is cleared because the copy is one concrete
+// instant, not a series that could be expanded again.
+func occurrence(e Event, start time.Time) Event {
+ o := e
+ o.RawRule = ""
+ o.Rule = nil
+ o.ExDates = nil
+ o.RecurrenceID = e.Start
+ o.Start = start
+ if d := e.Duration(); d > 0 {
+ o.End = start.Add(d)
+ } else {
+ o.End = time.Time{}
+ }
+ return o
+}
+
+// overrideIndex maps each series UID to the occurrence starts that an explicit
+// override event replaces, keyed by instant so an occurrence is matched however
+// its zone was written.
+func overrideIndex(events []Event) map[string]map[int64]bool {
+ out := map[string]map[int64]bool{}
+ for _, e := range events {
+ if e.RecurrenceID.IsZero() || e.UID == "" {
+ continue
+ }
+ if out[e.UID] == nil {
+ out[e.UID] = map[int64]bool{}
+ }
+ out[e.UID][e.RecurrenceID.UnixNano()] = true
+ }
+ return out
+}
+
+// excluded reports whether the start matches an EXDATE value.
+func excluded(exDates []time.Time, start time.Time) bool {
+ for _, x := range exDates {
+ if x.Equal(start) {
+ return true
+ }
+ }
+ return false
+}
+
+// within reports whether t falls inside the closed window, treating a zero
+// bound as open.
+func within(t, from, to time.Time) bool {
+ if !from.IsZero() && t.Before(from) {
+ return false
+ }
+ if !to.IsZero() && t.After(to) {
+ return false
+ }
+ return true
+}
diff --git a/ics/expand_test.go b/ics/expand_test.go
new file mode 100644
index 0000000..eb2b34a
--- /dev/null
+++ b/ics/expand_test.go
@@ -0,0 +1,163 @@
+package ics
+
+import (
+ "testing"
+ "time"
+
+ "github.com/google/go-cmp/cmp"
+ "github.com/google/go-cmp/cmp/cmpopts"
+)
+
+// weekly builds a recurring master event from an RRULE value.
+func weekly(t *testing.T, uid, summary, rule string, start time.Time, d time.Duration) Event {
+ t.Helper()
+ r, err := ParseRRULE(rule)
+ if err != nil {
+ t.Fatalf("ParseRRULE(%q): %v", rule, err)
+ }
+ e := Event{UID: uid, Summary: summary, Start: start, RawRule: rule, Rule: &r}
+ if d > 0 {
+ e.End = start.Add(d)
+ }
+ return e
+}
+
+// startStrings renders the start times of expanded events.
+func startStrings(events []Event) []string {
+ out := make([]string, len(events))
+ for i, e := range events {
+ out[i] = e.Start.Format("2006-01-02 15:04")
+ }
+ return out
+}
+
+func TestExpandSeriesProducesEveryOccurrence(t *testing.T) {
+ t.Parallel()
+ master := weekly(t, "standup@example", "Standup", "FREQ=WEEKLY;BYDAY=MO,WE", at(2024, time.March, 4, 9), 30*time.Minute)
+ got, report := Expand([]Event{master}, at(2024, time.March, 1, 0), at(2024, time.March, 15, 23))
+ want := []string{
+ "2024-03-04 09:00", "2024-03-06 09:00",
+ "2024-03-11 09:00", "2024-03-13 09:00",
+ }
+ if diff := cmp.Diff(want, startStrings(got)); diff != "" {
+ t.Errorf("occurrences mismatch (-want +got):\n%s", diff)
+ }
+ if report.Occurrences != 4 {
+ t.Errorf("want 4 occurrences reported, got %d", report.Occurrences)
+ }
+ for _, e := range got {
+ if e.Recurs() {
+ t.Errorf("occurrence %s still carries a rule", e.Start)
+ }
+ if e.UID != master.UID {
+ t.Errorf("occurrence %s lost the series UID", e.Start)
+ }
+ if e.End.Sub(e.Start) != 30*time.Minute {
+ t.Errorf("occurrence %s lost the series duration", e.Start)
+ }
+ }
+}
+
+func TestExpandSeriesStartingBeforeTheWindow(t *testing.T) {
+ t.Parallel()
+ // A weekly meeting running for years contributes its in-window occurrences
+ // even though the master event predates the window by a decade.
+ master := weekly(t, "oneone@example", "1:1", "FREQ=WEEKLY;BYDAY=TH", at(2014, time.January, 2, 15), time.Hour)
+ got, _ := Expand([]Event{master}, at(2024, time.March, 1, 0), at(2024, time.March, 31, 23))
+ want := []string{"2024-03-07 15:00", "2024-03-14 15:00", "2024-03-21 15:00", "2024-03-28 15:00"}
+ if diff := cmp.Diff(want, startStrings(got)); diff != "" {
+ t.Errorf("occurrences mismatch (-want +got):\n%s", diff)
+ }
+}
+
+func TestExpandHonorsExDate(t *testing.T) {
+ t.Parallel()
+ master := weekly(t, "standup@example", "Standup", "FREQ=WEEKLY;BYDAY=MO", at(2024, time.March, 4, 9), 0)
+ master.ExDates = []time.Time{at(2024, time.March, 11, 9)}
+ got, report := Expand([]Event{master}, at(2024, time.March, 1, 0), at(2024, time.March, 20, 23))
+ want := []string{"2024-03-04 09:00", "2024-03-18 09:00"}
+ if diff := cmp.Diff(want, startStrings(got)); diff != "" {
+ t.Errorf("occurrences mismatch (-want +got):\n%s", diff)
+ }
+ if report.Excluded != 1 {
+ t.Errorf("want 1 exclusion reported, got %d", report.Excluded)
+ }
+}
+
+func TestExpandOverrideReplacesGeneratedOccurrence(t *testing.T) {
+ t.Parallel()
+ master := weekly(t, "standup@example", "Standup", "FREQ=WEEKLY;BYDAY=MO", at(2024, time.March, 4, 9), 0)
+ // The calendar moved the second occurrence to the afternoon. Both the
+ // generated 09:00 and the moved 14:00 would otherwise land in the vault.
+ override := Event{
+ UID: "standup@example",
+ Summary: "Standup (moved)",
+ Start: at(2024, time.March, 11, 14),
+ RecurrenceID: at(2024, time.March, 11, 9),
+ }
+ got, report := Expand([]Event{master, override}, at(2024, time.March, 1, 0), at(2024, time.March, 20, 23))
+ want := []string{"2024-03-04 09:00", "2024-03-11 14:00", "2024-03-18 09:00"}
+ if diff := cmp.Diff(want, startStrings(got)); diff != "" {
+ t.Errorf("occurrences mismatch (-want +got):\n%s", diff)
+ }
+ if report.Overridden != 1 {
+ t.Errorf("want 1 override reported, got %d", report.Overridden)
+ }
+}
+
+func TestExpandUnexpandableRuleKeepsFirstOccurrence(t *testing.T) {
+ t.Parallel()
+ // A frequency midden does not expand still contributes the event as written,
+ // and is counted so the caller can say the series is incomplete.
+ master := Event{UID: "odd@example", Summary: "Hourly", Start: at(2024, time.March, 4, 9), RawRule: "FREQ=HOURLY"}
+ got, report := Expand([]Event{master}, at(2024, time.March, 1, 0), at(2024, time.March, 20, 23))
+ if diff := cmp.Diff([]string{"2024-03-04 09:00"}, startStrings(got)); diff != "" {
+ t.Errorf("occurrences mismatch (-want +got):\n%s", diff)
+ }
+ if report.Unexpanded != 1 {
+ t.Errorf("want 1 unexpanded series reported, got %d", report.Unexpanded)
+ }
+}
+
+func TestExpandFiltersPlainEventsToWindow(t *testing.T) {
+ t.Parallel()
+ events := []Event{
+ {UID: "a", Summary: "before", Start: at(2024, time.February, 1, 9)},
+ {UID: "b", Summary: "inside", Start: at(2024, time.March, 5, 9)},
+ {UID: "c", Summary: "after", Start: at(2024, time.April, 1, 9)},
+ }
+ got, report := Expand(events, at(2024, time.March, 1, 0), at(2024, time.March, 31, 23))
+ if diff := cmp.Diff([]string{"2024-03-05 09:00"}, startStrings(got)); diff != "" {
+ t.Errorf("occurrences mismatch (-want +got):\n%s", diff)
+ }
+ if diff := cmp.Diff(ExpandReport{}, report, cmpopts.EquateEmpty()); diff != "" {
+ t.Errorf("want an empty report for plain events (-want +got):\n%s", diff)
+ }
+}
+
+func TestExpandOrdersMixedSourcesByStart(t *testing.T) {
+ t.Parallel()
+ events := []Event{
+ {UID: "plain", Summary: "one off", Start: at(2024, time.March, 12, 8)},
+ weekly(t, "s", "Standup", "FREQ=WEEKLY;BYDAY=MO", at(2024, time.March, 4, 9), 0),
+ }
+ got, _ := Expand(events, at(2024, time.March, 1, 0), at(2024, time.March, 20, 23))
+ want := []string{"2024-03-04 09:00", "2024-03-11 09:00", "2024-03-12 08:00", "2024-03-18 09:00"}
+ if diff := cmp.Diff(want, startStrings(got)); diff != "" {
+ t.Errorf("ordering mismatch (-want +got):\n%s", diff)
+ }
+}
+
+func TestExpandReportsTruncatedUnboundedSeries(t *testing.T) {
+ t.Parallel()
+ master := weekly(t, "forever@example", "Forever", "FREQ=DAILY", at(2024, time.March, 4, 9), 0)
+ // With no window end and no COUNT or UNTIL there is nothing to stop the
+ // walk, so the series is reported rather than expanded.
+ got, report := Expand([]Event{master}, time.Time{}, time.Time{})
+ if len(got) != 0 {
+ t.Errorf("want no occurrences from an unbounded walk, got %d", len(got))
+ }
+ if report.Truncated != 1 {
+ t.Errorf("want 1 truncated series reported, got %d", report.Truncated)
+ }
+}
diff --git a/ics/ics.go b/ics/ics.go
index 6b500f3..19bc45a 100644
--- a/ics/ics.go
+++ b/ics/ics.go
@@ -1,12 +1,14 @@
// Package ics parses the small iCalendar subset midden ingests from local .ics exports.
//
// The parser handles line folding, VEVENT records, and the SUMMARY, DTSTART,
-// DTEND, LOCATION, DESCRIPTION, UID, and RRULE properties. Components nested
-// inside a VEVENT (VALARM and friends) are skipped so their properties never
-// touch the parent event. TZID parameters resolve through time.LoadLocation
-// with a fallback to the machine's local zone; the UTC suffix Z is honored.
-// Parsed times are anchored to the local zone. Events whose DTSTART is
-// missing or unparseable are dropped and counted rather than returned.
+// DTEND, LOCATION, DESCRIPTION, UID, RRULE, EXDATE, and RECURRENCE-ID
+// properties. Components nested inside a VEVENT (VALARM and friends) are
+// skipped so their properties never touch the parent event. TZID parameters
+// resolve through time.LoadLocation with a fallback to the machine's local
+// zone; the UTC suffix Z is honored. Parsed times are anchored to the local
+// zone. Events whose DTSTART is missing or unparseable are dropped and counted
+// rather than returned. Parse returns series as written; call Expand to turn a
+// recurring series into the occurrences it actually produced.
package ics
import (
@@ -33,8 +35,30 @@ type Event struct {
Description string
// AllDay reports whether DTSTART carried a date-only value.
AllDay bool
- // Recurs reports whether the event carries an RRULE. Recurrences are not expanded.
- Recurs bool
+ // RawRule is the verbatim RRULE value, or empty when the event does not recur.
+ RawRule string
+ // Rule is the parsed RRULE, or nil when the event does not recur or its rule
+ // uses a frequency midden does not expand. A non-empty RawRule with a nil
+ // Rule is a series that will not be expanded.
+ Rule *Recurrence
+ // ExDates are the occurrence start times EXDATE removes from the series.
+ ExDates []time.Time
+ // RecurrenceID identifies which occurrence of a series this event replaces,
+ // or the zero time when the event is not an override.
+ RecurrenceID time.Time
+}
+
+// Recurs reports whether the event carries a recurrence rule.
+func (e Event) Recurs() bool {
+ return e.RawRule != ""
+}
+
+// Duration returns how long the event lasts, or zero when DTEND is absent.
+func (e Event) Duration() time.Duration {
+ if e.End.IsZero() || e.End.Before(e.Start) {
+ return 0
+ }
+ return e.End.Sub(e.Start)
}
// Parse reads an iCalendar stream and returns every VEVENT it contains plus
@@ -125,7 +149,20 @@ func applyProperty(e *Event, line string) {
case "UID":
e.UID = unescape(value)
case "RRULE":
- e.Recurs = true
+ e.RawRule = value
+ if rule, err := ParseRRULE(value); err == nil {
+ e.Rule = &rule
+ }
+ case "EXDATE":
+ for v := range strings.SplitSeq(value, ",") {
+ if t, _, ok := parseTime(strings.TrimSpace(v), params); ok {
+ e.ExDates = append(e.ExDates, t)
+ }
+ }
+ case "RECURRENCE-ID":
+ if t, _, ok := parseTime(value, params); ok {
+ e.RecurrenceID = t
+ }
case "DTSTART":
if t, allDay, ok := parseTime(value, params); ok {
e.Start = t
diff --git a/ics/ics_test.go b/ics/ics_test.go
index fd244ea..0cfdfb0 100644
--- a/ics/ics_test.go
+++ b/ics/ics_test.go
@@ -158,7 +158,7 @@ END:VEVENT
Summary: "Timed",
Start: time.Date(2026, 6, 16, 14, 0, 0, 0, time.UTC),
}},
- }, { // Test 9: An RRULE property sets Recurs without expanding occurrences.
+ }, { // Test 9: An RRULE property is kept verbatim and parsed for later expansion.
In: strings.NewReader(`BEGIN:VEVENT
SUMMARY:Standup
DTSTART:20260616T140000Z
@@ -168,7 +168,13 @@ END:VEVENT
WantEvents: []Event{{
Summary: "Standup",
Start: time.Date(2026, 6, 16, 14, 0, 0, 0, time.UTC),
- Recurs: true,
+ RawRule: "FREQ=WEEKLY;BYDAY=MO",
+ Rule: &Recurrence{
+ Freq: Weekly,
+ Interval: 1,
+ WeekStart: time.Monday,
+ ByDay: []WeekDayNum{{Day: time.Monday}},
+ },
}},
}, { // Test 10: A malformed DTSTART drops the event and counts it as skipped.
In: strings.NewReader(`BEGIN:VEVENT
diff --git a/ics/rrule.go b/ics/rrule.go
new file mode 100644
index 0000000..d29511a
--- /dev/null
+++ b/ics/rrule.go
@@ -0,0 +1,510 @@
+package ics
+
+import (
+ "fmt"
+ "sort"
+ "strconv"
+ "strings"
+ "time"
+)
+
+// maxPeriods bounds how many periods a single rule is walked before the
+// expansion gives up. A rule with no COUNT or UNTIL is bounded only by the
+// requested window, so a pathological one would otherwise walk forever.
+const maxPeriods = 100000
+
+// Frequency is the RRULE FREQ value. Only the frequencies personal calendars
+// actually use are expanded; anything else is left unexpanded rather than
+// guessed at.
+type Frequency string
+
+// The expanded frequencies.
+const (
+ Daily Frequency = "DAILY"
+ Weekly Frequency = "WEEKLY"
+ Monthly Frequency = "MONTHLY"
+ Yearly Frequency = "YEARLY"
+)
+
+// WeekDayNum is one BYDAY entry: a weekday with an optional ordinal that
+// selects, for example, the second Tuesday or the last Friday of a month.
+type WeekDayNum struct {
+ // Day is the weekday the entry selects.
+ Day time.Weekday
+ // Ordinal positions the weekday within the period, counting from the end
+ // when negative. Zero means every matching weekday.
+ Ordinal int
+}
+
+// Recurrence is the RRULE subset midden expands. BYSETPOS, BYYEARDAY, and
+// BYWEEKNO are not supported; a rule using them still expands on its remaining
+// parts rather than being dropped, so the result may hold more occurrences than
+// the calendar shows.
+type Recurrence struct {
+ // Freq is the base frequency of the rule.
+ Freq Frequency
+ // Interval is the number of periods between occurrences, at least one.
+ Interval int
+ // Count caps the number of occurrences generated from DTSTART, or zero for
+ // no cap.
+ Count int
+ // Until is the inclusive last instant an occurrence may fall on, or the
+ // zero time for no bound.
+ Until time.Time
+ // ByDay restricts or selects weekdays within each period.
+ ByDay []WeekDayNum
+ // ByMonthDay restricts days of the month, counting from the end when negative.
+ ByMonthDay []int
+ // ByMonth restricts which calendar months may hold occurrences.
+ ByMonth []time.Month
+ // WeekStart is the first day of the week, which decides week boundaries for
+ // a weekly rule with an interval above one.
+ WeekStart time.Weekday
+}
+
+// ParseRRULE parses an RRULE property value. An unrecognized or absent FREQ is
+// an error, because expanding a rule whose base frequency is unknown would
+// invent occurrences rather than omit them.
+func ParseRRULE(value string) (Recurrence, error) {
+ r := Recurrence{Interval: 1, WeekStart: time.Monday}
+ for part := range strings.SplitSeq(value, ";") {
+ key, val, ok := strings.Cut(part, "=")
+ if !ok || val == "" {
+ continue
+ }
+ switch strings.ToUpper(strings.TrimSpace(key)) {
+ case "FREQ":
+ r.Freq = Frequency(strings.ToUpper(val))
+ case "INTERVAL":
+ n, err := strconv.Atoi(val)
+ if err != nil || n < 1 {
+ return Recurrence{}, fmt.Errorf("rrule interval %q: not a positive number", val)
+ }
+ r.Interval = n
+ case "COUNT":
+ n, err := strconv.Atoi(val)
+ if err != nil || n < 1 {
+ return Recurrence{}, fmt.Errorf("rrule count %q: not a positive number", val)
+ }
+ r.Count = n
+ case "UNTIL":
+ t, _, ok := parseTime(val, "")
+ if !ok {
+ return Recurrence{}, fmt.Errorf("rrule until %q: unparseable time", val)
+ }
+ r.Until = t
+ case "BYDAY":
+ days, err := parseByDay(val)
+ if err != nil {
+ return Recurrence{}, err
+ }
+ r.ByDay = days
+ case "BYMONTHDAY":
+ days, err := parseInts(val, "bymonthday")
+ if err != nil {
+ return Recurrence{}, err
+ }
+ r.ByMonthDay = days
+ case "BYMONTH":
+ months, err := parseInts(val, "bymonth")
+ if err != nil {
+ return Recurrence{}, err
+ }
+ for _, m := range months {
+ if m < 1 || m > 12 {
+ return Recurrence{}, fmt.Errorf("rrule bymonth %d: out of range", m)
+ }
+ r.ByMonth = append(r.ByMonth, time.Month(m))
+ }
+ case "WKST":
+ d, ok := parseWeekday(val)
+ if !ok {
+ return Recurrence{}, fmt.Errorf("rrule wkst %q: not a weekday", val)
+ }
+ r.WeekStart = d
+ }
+ }
+ switch r.Freq {
+ case Daily, Weekly, Monthly, Yearly:
+ return r, nil
+ case "":
+ return Recurrence{}, fmt.Errorf("rrule has no freq")
+ }
+ return Recurrence{}, fmt.Errorf("rrule freq %q: not expanded", r.Freq)
+}
+
+// Occurrences returns every start time the rule generates for dtstart that
+// falls inside the window, ordered ascending, and reports whether the walk was
+// cut short by the period cap. Generation always begins at dtstart so COUNT is
+// applied to the real series rather than to the part of it inside the window. A
+// zero to leaves the window unbounded on the far side, which only terminates
+// when the rule itself does.
+func (r Recurrence) Occurrences(dtstart, from, to time.Time) ([]time.Time, bool) {
+ interval := r.Interval
+ if interval < 1 {
+ interval = 1
+ }
+ end := to
+ if !r.Until.IsZero() && (end.IsZero() || r.Until.Before(end)) {
+ end = r.Until
+ }
+ if end.IsZero() && r.Count == 0 {
+ // Nothing bounds the walk, so refuse rather than spin.
+ return nil, true
+ }
+
+ var out []time.Time
+ generated := 0
+ period := r.periodAnchor(dtstart)
+ for guard := 0; ; guard++ {
+ if guard >= maxPeriods {
+ return out, true
+ }
+ if !end.IsZero() && period.After(end) {
+ break
+ }
+ for _, c := range r.candidates(period, dtstart) {
+ if c.Before(dtstart) {
+ continue
+ }
+ if !end.IsZero() && c.After(end) {
+ continue
+ }
+ generated++
+ if r.Count > 0 && generated > r.Count {
+ return out, false
+ }
+ if from.IsZero() || !c.Before(from) {
+ out = append(out, c)
+ }
+ }
+ if r.Count > 0 && generated >= r.Count {
+ break
+ }
+ period = r.advance(period, interval)
+ }
+ return out, false
+}
+
+// periodAnchor normalizes dtstart to the start of the period containing it, so
+// stepping by interval never has to clamp a day of month that the next period
+// does not have.
+func (r Recurrence) periodAnchor(dtstart time.Time) time.Time {
+ y, m, d := dtstart.Date()
+ loc := dtstart.Location()
+ switch r.Freq {
+ case Weekly:
+ day := time.Date(y, m, d, 0, 0, 0, 0, loc)
+ back := (int(day.Weekday()) - int(r.WeekStart) + 7) % 7
+ return day.AddDate(0, 0, -back)
+ case Monthly:
+ return time.Date(y, m, 1, 0, 0, 0, 0, loc)
+ case Yearly:
+ return time.Date(y, time.January, 1, 0, 0, 0, 0, loc)
+ default:
+ return time.Date(y, m, d, 0, 0, 0, 0, loc)
+ }
+}
+
+// advance steps the period anchor forward by interval periods.
+func (r Recurrence) advance(period time.Time, interval int) time.Time {
+ switch r.Freq {
+ case Weekly:
+ return period.AddDate(0, 0, 7*interval)
+ case Monthly:
+ return addMonths(period, interval)
+ case Yearly:
+ return time.Date(period.Year()+interval, time.January, 1, 0, 0, 0, 0, period.Location())
+ default:
+ return period.AddDate(0, 0, interval)
+ }
+}
+
+// candidates returns the occurrence times the rule produces inside the period
+// beginning at the anchor, ordered ascending and carrying dtstart's clock time.
+func (r Recurrence) candidates(period, dtstart time.Time) []time.Time {
+ var days []time.Time
+ switch r.Freq {
+ case Daily:
+ days = []time.Time{period}
+ case Weekly:
+ days = r.weeklyDays(period, dtstart)
+ case Monthly:
+ days = r.monthDays(period.Year(), period.Month(), dtstart)
+ case Yearly:
+ days = r.yearlyDays(period.Year(), dtstart)
+ }
+ out := make([]time.Time, 0, len(days))
+ for _, d := range days {
+ if !r.allows(d) {
+ continue
+ }
+ out = append(out, withClock(d, dtstart))
+ }
+ sort.Slice(out, func(i, j int) bool { return out[i].Before(out[j]) })
+ return out
+}
+
+// weeklyDays returns the days a weekly rule selects inside the week beginning
+// at the anchor, defaulting to dtstart's own weekday when BYDAY is absent.
+func (r Recurrence) weeklyDays(period, dtstart time.Time) []time.Time {
+ wanted := map[time.Weekday]bool{}
+ if len(r.ByDay) == 0 {
+ wanted[dtstart.Weekday()] = true
+ }
+ for _, wd := range r.ByDay {
+ wanted[wd.Day] = true
+ }
+ var out []time.Time
+ for i := range 7 {
+ d := period.AddDate(0, 0, i)
+ if wanted[d.Weekday()] {
+ out = append(out, d)
+ }
+ }
+ return out
+}
+
+// monthDays returns the days a monthly or yearly rule selects inside one month.
+// BYDAY ordinals count from the end of the month when negative. Giving both
+// BYDAY and BYMONTHDAY narrows to the days satisfying both, which is what makes
+// FREQ=MONTHLY;BYDAY=FR;BYMONTHDAY=13 mean Friday the thirteenth rather than
+// every Friday plus every thirteenth. With neither part the rule falls back to
+// dtstart's day of month, which months too short to hold it simply skip.
+func (r Recurrence) monthDays(year int, month time.Month, dtstart time.Time) []time.Time {
+ loc := dtstart.Location()
+ last := daysInMonth(year, month)
+ var selected map[int]bool
+ switch {
+ case len(r.ByMonthDay) > 0 && len(r.ByDay) > 0:
+ selected = intersectDays(r.monthDayNumbers(last), r.weekdayNumbers(year, month, loc, last))
+ case len(r.ByMonthDay) > 0:
+ selected = r.monthDayNumbers(last)
+ case len(r.ByDay) > 0:
+ selected = r.weekdayNumbers(year, month, loc, last)
+ default:
+ selected = map[int]bool{dtstart.Day(): true}
+ }
+ out := make([]time.Time, 0, len(selected))
+ for day := 1; day <= last; day++ {
+ if selected[day] {
+ out = append(out, time.Date(year, month, day, 0, 0, 0, 0, loc))
+ }
+ }
+ return out
+}
+
+// monthDayNumbers returns the days of month BYMONTHDAY selects, resolving
+// negative values against the length of the month.
+func (r Recurrence) monthDayNumbers(last int) map[int]bool {
+ out := map[int]bool{}
+ for _, d := range r.ByMonthDay {
+ if d < 0 {
+ d = last + 1 + d
+ }
+ if d >= 1 && d <= last {
+ out[d] = true
+ }
+ }
+ return out
+}
+
+// weekdayNumbers returns the days of month BYDAY selects, resolving ordinals
+// against the weekdays present in that month.
+func (r Recurrence) weekdayNumbers(year int, month time.Month, loc *time.Location, last int) map[int]bool {
+ out := map[int]bool{}
+ for _, wd := range r.ByDay {
+ if wd.Ordinal != 0 {
+ if day := nthWeekday(year, month, wd.Day, wd.Ordinal, loc); day != 0 {
+ out[day] = true
+ }
+ continue
+ }
+ for day := 1; day <= last; day++ {
+ if time.Date(year, month, day, 0, 0, 0, 0, loc).Weekday() == wd.Day {
+ out[day] = true
+ }
+ }
+ }
+ return out
+}
+
+// intersectDays returns the days present in both sets.
+func intersectDays(a, b map[int]bool) map[int]bool {
+ out := map[int]bool{}
+ for day := range a {
+ if b[day] {
+ out[day] = true
+ }
+ }
+ return out
+}
+
+// yearlyDays returns the days a yearly rule selects inside one year, defaulting
+// to dtstart's own month when BYMONTH is absent.
+func (r Recurrence) yearlyDays(year int, dtstart time.Time) []time.Time {
+ months := r.ByMonth
+ if len(months) == 0 {
+ months = []time.Month{dtstart.Month()}
+ }
+ var out []time.Time
+ for _, m := range months {
+ out = append(out, r.monthDays(year, m, dtstart)...)
+ }
+ return out
+}
+
+// allows applies the BY parts that act as filters rather than as generators for
+// the rule's frequency. A weekly rule already generated its weekdays from
+// BYDAY, and a monthly or yearly rule already generated its days, so applying
+// those parts again here would be a no-op at best.
+func (r Recurrence) allows(day time.Time) bool {
+ if len(r.ByMonth) > 0 && !containsMonth(r.ByMonth, day.Month()) {
+ return false
+ }
+ if r.Freq != Daily {
+ return true
+ }
+ if len(r.ByDay) > 0 && !containsWeekday(r.ByDay, day.Weekday()) {
+ return false
+ }
+ if len(r.ByMonthDay) > 0 {
+ last := daysInMonth(day.Year(), day.Month())
+ if !containsInt(r.ByMonthDay, day.Day()) && !containsInt(r.ByMonthDay, day.Day()-last-1) {
+ return false
+ }
+ }
+ return true
+}
+
+// withClock returns the day carrying dtstart's wall-clock time. Wall clock is
+// preserved rather than elapsed time so an event stays at the hour the calendar
+// shows it across a daylight-saving shift.
+func withClock(day, dtstart time.Time) time.Time {
+ h, m, s := dtstart.Clock()
+ y, mo, d := day.Date()
+ return time.Date(y, mo, d, h, m, s, 0, dtstart.Location())
+}
+
+// nthWeekday returns the day of month of the ordinal-th given weekday, counting
+// from the end of the month when the ordinal is negative, or zero when the
+// month has no such day.
+func nthWeekday(year int, month time.Month, want time.Weekday, ordinal int, loc *time.Location) int {
+ last := daysInMonth(year, month)
+ var days []int
+ for day := 1; day <= last; day++ {
+ if time.Date(year, month, day, 0, 0, 0, 0, loc).Weekday() == want {
+ days = append(days, day)
+ }
+ }
+ if ordinal > 0 && ordinal <= len(days) {
+ return days[ordinal-1]
+ }
+ if ordinal < 0 && -ordinal <= len(days) {
+ return days[len(days)+ordinal]
+ }
+ return 0
+}
+
+// addMonths shifts a first-of-month anchor forward by delta months.
+func addMonths(t time.Time, delta int) time.Time {
+ total := int(t.Month()) - 1 + delta
+ year := t.Year() + total/12
+ month := time.Month(total%12 + 1)
+ return time.Date(year, month, 1, 0, 0, 0, 0, t.Location())
+}
+
+// daysInMonth returns the number of days in the given month.
+func daysInMonth(year int, month time.Month) int {
+ return time.Date(year, month+1, 0, 0, 0, 0, 0, time.UTC).Day()
+}
+
+// parseByDay parses a BYDAY value such as "MO,WE,FR" or "-1FR,2TU".
+func parseByDay(value string) ([]WeekDayNum, error) {
+ var out []WeekDayNum
+ for part := range strings.SplitSeq(value, ",") {
+ part = strings.TrimSpace(part)
+ if len(part) < 2 {
+ return nil, fmt.Errorf("rrule byday %q: too short", part)
+ }
+ name := part[len(part)-2:]
+ day, ok := parseWeekday(name)
+ if !ok {
+ return nil, fmt.Errorf("rrule byday %q: not a weekday", part)
+ }
+ entry := WeekDayNum{Day: day}
+ if prefix := part[:len(part)-2]; prefix != "" {
+ n, err := strconv.Atoi(prefix)
+ if err != nil || n == 0 {
+ return nil, fmt.Errorf("rrule byday %q: bad ordinal", part)
+ }
+ entry.Ordinal = n
+ }
+ out = append(out, entry)
+ }
+ return out, nil
+}
+
+// parseInts parses a comma-separated list of signed integers.
+func parseInts(value, label string) ([]int, error) {
+ var out []int
+ for part := range strings.SplitSeq(value, ",") {
+ n, err := strconv.Atoi(strings.TrimSpace(part))
+ if err != nil {
+ return nil, fmt.Errorf("rrule %s %q: not a number", label, part)
+ }
+ out = append(out, n)
+ }
+ return out, nil
+}
+
+// parseWeekday converts a two-letter iCalendar weekday abbreviation.
+func parseWeekday(s string) (time.Weekday, bool) {
+ switch strings.ToUpper(strings.TrimSpace(s)) {
+ case "SU":
+ return time.Sunday, true
+ case "MO":
+ return time.Monday, true
+ case "TU":
+ return time.Tuesday, true
+ case "WE":
+ return time.Wednesday, true
+ case "TH":
+ return time.Thursday, true
+ case "FR":
+ return time.Friday, true
+ case "SA":
+ return time.Saturday, true
+ }
+ return 0, false
+}
+
+// containsMonth reports whether the month appears in the list.
+func containsMonth(months []time.Month, m time.Month) bool {
+ for _, x := range months {
+ if x == m {
+ return true
+ }
+ }
+ return false
+}
+
+// containsWeekday reports whether the weekday appears in the BYDAY list.
+func containsWeekday(days []WeekDayNum, d time.Weekday) bool {
+ for _, x := range days {
+ if x.Day == d {
+ return true
+ }
+ }
+ return false
+}
+
+// containsInt reports whether n appears in the list.
+func containsInt(list []int, n int) bool {
+ for _, x := range list {
+ if x == n {
+ return true
+ }
+ }
+ return false
+}
diff --git a/ics/rrule_test.go b/ics/rrule_test.go
new file mode 100644
index 0000000..932df3a
--- /dev/null
+++ b/ics/rrule_test.go
@@ -0,0 +1,236 @@
+package ics
+
+import (
+ "fmt"
+ "testing"
+ "time"
+
+ "github.com/google/go-cmp/cmp"
+ "github.com/google/go-cmp/cmp/cmpopts"
+)
+
+// at builds a local timestamp for the given calendar date and hour.
+func at(year int, month time.Month, day, hour int) time.Time {
+ return time.Date(year, month, day, hour, 0, 0, 0, time.Local)
+}
+
+// occurrenceStrings renders occurrence times for comparison.
+func occurrenceStrings(times []time.Time) []string {
+ out := make([]string, len(times))
+ for i, t := range times {
+ out[i] = t.Format("2006-01-02 15:04")
+ }
+ return out
+}
+
+func TestParseRRULE(t *testing.T) {
+ t.Parallel()
+ tests := []struct {
+ Rule string
+ WantRule Recurrence
+ Want bool
+ }{{ // Test 0: A plain weekly rule takes the default interval and week start.
+ Rule: "FREQ=WEEKLY",
+ WantRule: Recurrence{Freq: Weekly, Interval: 1, WeekStart: time.Monday},
+ }, { // Test 1: Every documented part is parsed.
+ Rule: "FREQ=MONTHLY;INTERVAL=2;COUNT=5;BYDAY=-1FR,TU;BYMONTHDAY=13,-1;BYMONTH=3,6;WKST=SU",
+ WantRule: Recurrence{
+ Freq: Monthly, Interval: 2, Count: 5,
+ ByDay: []WeekDayNum{{Day: time.Friday, Ordinal: -1}, {Day: time.Tuesday}},
+ ByMonthDay: []int{13, -1},
+ ByMonth: []time.Month{time.March, time.June},
+ WeekStart: time.Sunday,
+ },
+ }, { // Test 2: UNTIL in UTC resolves to a local instant.
+ Rule: "FREQ=DAILY;UNTIL=20240310T235959Z",
+ WantRule: Recurrence{Freq: Daily, Interval: 1, WeekStart: time.Monday, Until: time.Date(2024, time.March, 10, 23, 59, 59, 0, time.UTC).Local()},
+ }, { // Test 3: A missing FREQ is rejected rather than defaulted.
+ Rule: "INTERVAL=2",
+ Want: true,
+ }, { // Test 4: An unsupported FREQ is rejected rather than silently expanded wrong.
+ Rule: "FREQ=SECONDLY",
+ Want: true,
+ }, { // Test 5: A zero interval is rejected.
+ Rule: "FREQ=DAILY;INTERVAL=0",
+ Want: true,
+ }, { // Test 6: A bad weekday is rejected.
+ Rule: "FREQ=WEEKLY;BYDAY=XX",
+ Want: true,
+ }, { // Test 7: A zero BYDAY ordinal is rejected.
+ Rule: "FREQ=MONTHLY;BYDAY=0FR",
+ Want: true,
+ }, { // Test 8: An out-of-range BYMONTH is rejected.
+ Rule: "FREQ=YEARLY;BYMONTH=13",
+ Want: true,
+ }}
+ for testNum, test := range tests {
+ t.Run(fmt.Sprintf("test %d", testNum), func(t *testing.T) {
+ t.Parallel()
+ got, err := ParseRRULE(test.Rule)
+ if test.Want {
+ if err == nil {
+ t.Fatalf("want error, got %+v", got)
+ }
+ return
+ }
+ if err != nil {
+ t.Fatalf("ParseRRULE: %v", err)
+ }
+ if diff := cmp.Diff(test.WantRule, got, cmpopts.EquateEmpty()); diff != "" {
+ t.Errorf("mismatch (-want +got):\n%s", diff)
+ }
+ })
+ }
+}
+
+func TestOccurrences(t *testing.T) {
+ t.Parallel()
+ tests := []struct {
+ DTStart time.Time
+ From time.Time
+ To time.Time
+ Rule string
+ Want []string
+ }{{ // Test 0: A weekly standup lands on each listed weekday.
+ Rule: "FREQ=WEEKLY;BYDAY=MO,WE,FR",
+ DTStart: at(2024, time.March, 4, 9), // a Monday
+ To: at(2024, time.March, 15, 23),
+ Want: []string{
+ "2024-03-04 09:00", "2024-03-06 09:00", "2024-03-08 09:00",
+ "2024-03-11 09:00", "2024-03-13 09:00", "2024-03-15 09:00",
+ },
+ }, { // Test 1: A weekly rule with no BYDAY repeats on dtstart's own weekday.
+ Rule: "FREQ=WEEKLY",
+ DTStart: at(2024, time.March, 5, 14), // a Tuesday
+ To: at(2024, time.March, 26, 23),
+ Want: []string{"2024-03-05 14:00", "2024-03-12 14:00", "2024-03-19 14:00", "2024-03-26 14:00"},
+ }, { // Test 2: A fortnightly one-to-one skips the intervening week.
+ Rule: "FREQ=WEEKLY;INTERVAL=2;BYDAY=TH",
+ DTStart: at(2024, time.March, 7, 15), // a Thursday
+ To: at(2024, time.April, 18, 23),
+ Want: []string{"2024-03-07 15:00", "2024-03-21 15:00", "2024-04-04 15:00", "2024-04-18 15:00"},
+ }, { // Test 3: A monthly rule repeats on dtstart's day of month.
+ Rule: "FREQ=MONTHLY",
+ DTStart: at(2024, time.January, 15, 10),
+ To: at(2024, time.April, 30, 23),
+ Want: []string{"2024-01-15 10:00", "2024-02-15 10:00", "2024-03-15 10:00", "2024-04-15 10:00"},
+ }, { // Test 4: A monthly rule on the 31st skips months too short to hold it.
+ Rule: "FREQ=MONTHLY;BYMONTHDAY=31",
+ DTStart: at(2024, time.January, 31, 8),
+ To: at(2024, time.June, 30, 23),
+ Want: []string{"2024-01-31 08:00", "2024-03-31 08:00", "2024-05-31 08:00"},
+ }, { // Test 5: An ordinal BYDAY selects the second Tuesday of each month.
+ Rule: "FREQ=MONTHLY;BYDAY=2TU",
+ DTStart: at(2024, time.January, 9, 18),
+ To: at(2024, time.April, 30, 23),
+ Want: []string{"2024-01-09 18:00", "2024-02-13 18:00", "2024-03-12 18:00", "2024-04-09 18:00"},
+ }, { // Test 6: A negative ordinal selects the last Friday of each month.
+ Rule: "FREQ=MONTHLY;BYDAY=-1FR",
+ DTStart: at(2024, time.January, 26, 17),
+ To: at(2024, time.March, 31, 23),
+ Want: []string{"2024-01-26 17:00", "2024-02-23 17:00", "2024-03-29 17:00"},
+ }, { // Test 7: BYDAY and BYMONTHDAY together narrow to Friday the thirteenth.
+ Rule: "FREQ=MONTHLY;BYDAY=FR;BYMONTHDAY=13",
+ DTStart: at(2024, time.September, 13, 12),
+ To: at(2025, time.December, 31, 23),
+ Want: []string{"2024-09-13 12:00", "2024-12-13 12:00", "2025-06-13 12:00"},
+ }, { // Test 8: A yearly rule is the birthday case.
+ Rule: "FREQ=YEARLY",
+ DTStart: at(2020, time.July, 4, 0),
+ To: at(2023, time.December, 31, 23),
+ Want: []string{"2020-07-04 00:00", "2021-07-04 00:00", "2022-07-04 00:00", "2023-07-04 00:00"},
+ }, { // Test 9: A yearly rule with BYMONTH and an ordinal BYDAY tracks a moving holiday.
+ Rule: "FREQ=YEARLY;BYMONTH=5;BYDAY=2SU",
+ DTStart: at(2024, time.May, 12, 11),
+ To: at(2026, time.December, 31, 23),
+ Want: []string{"2024-05-12 11:00", "2025-05-11 11:00", "2026-05-10 11:00"},
+ }, { // Test 10: COUNT caps the series regardless of how far the window reaches.
+ Rule: "FREQ=DAILY;COUNT=3",
+ DTStart: at(2024, time.March, 1, 7),
+ To: at(2024, time.December, 31, 23),
+ Want: []string{"2024-03-01 07:00", "2024-03-02 07:00", "2024-03-03 07:00"},
+ }, { // Test 11: UNTIL ends the series before the window does.
+ Rule: "FREQ=DAILY;UNTIL=20240304T000000",
+ DTStart: at(2024, time.March, 1, 0),
+ To: at(2024, time.December, 31, 23),
+ Want: []string{"2024-03-01 00:00", "2024-03-02 00:00", "2024-03-03 00:00", "2024-03-04 00:00"},
+ }, { // Test 12: An interval above one on a daily rule skips days.
+ Rule: "FREQ=DAILY;INTERVAL=3",
+ DTStart: at(2024, time.March, 1, 6),
+ To: at(2024, time.March, 10, 23),
+ Want: []string{"2024-03-01 06:00", "2024-03-04 06:00", "2024-03-07 06:00", "2024-03-10 06:00"},
+ }, { // Test 13: COUNT is measured from dtstart, not from the start of the window.
+ Rule: "FREQ=DAILY;COUNT=5",
+ DTStart: at(2024, time.March, 1, 9),
+ From: at(2024, time.March, 4, 0),
+ To: at(2024, time.December, 31, 23),
+ Want: []string{"2024-03-04 09:00", "2024-03-05 09:00"},
+ }, { // Test 14: A window opening after the series ends yields nothing.
+ Rule: "FREQ=WEEKLY;COUNT=2",
+ DTStart: at(2024, time.March, 4, 9),
+ From: at(2025, time.January, 1, 0),
+ To: at(2025, time.December, 31, 23),
+ Want: nil,
+ }, { // Test 15: A daily rule filtered by BYDAY keeps only weekdays.
+ Rule: "FREQ=DAILY;BYDAY=SA,SU",
+ DTStart: at(2024, time.March, 1, 8), // a Friday
+ To: at(2024, time.March, 11, 23),
+ Want: []string{"2024-03-02 08:00", "2024-03-03 08:00", "2024-03-09 08:00", "2024-03-10 08:00"},
+ }}
+ for testNum, test := range tests {
+ t.Run(fmt.Sprintf("test %d", testNum), func(t *testing.T) {
+ t.Parallel()
+ rule, err := ParseRRULE(test.Rule)
+ if err != nil {
+ t.Fatalf("ParseRRULE(%q): %v", test.Rule, err)
+ }
+ got, truncated := rule.Occurrences(test.DTStart, test.From, test.To)
+ if truncated {
+ t.Fatalf("expansion was truncated")
+ }
+ if diff := cmp.Diff(test.Want, occurrenceStrings(got), cmpopts.EquateEmpty()); diff != "" {
+ t.Errorf("mismatch (-want +got):\n%s", diff)
+ }
+ })
+ }
+}
+
+func TestOccurrencesRefusesUnboundedWalk(t *testing.T) {
+ t.Parallel()
+ rule, err := ParseRRULE("FREQ=DAILY")
+ if err != nil {
+ t.Fatalf("ParseRRULE: %v", err)
+ }
+ // No COUNT, no UNTIL, and no window end leaves nothing to stop the walk.
+ got, truncated := rule.Occurrences(at(2024, time.March, 1, 9), time.Time{}, time.Time{})
+ if !truncated {
+ t.Errorf("want the walk refused, got %d occurrences", len(got))
+ }
+}
+
+func TestOccurrencesPreservesWallClockAcrossDST(t *testing.T) {
+ t.Parallel()
+ loc, err := time.LoadLocation("America/Chicago")
+ if err != nil {
+ t.Skipf("tzdata unavailable: %v", err)
+ }
+ rule, err := ParseRRULE("FREQ=WEEKLY;BYDAY=SU")
+ if err != nil {
+ t.Fatalf("ParseRRULE: %v", err)
+ }
+ // US daylight saving began on 2024-03-10. A weekly 09:00 meeting stays at
+ // 09:00 local rather than sliding to 08:00 or 10:00.
+ start := time.Date(2024, time.March, 3, 9, 0, 0, 0, loc)
+ got, truncated := rule.Occurrences(start, time.Time{}, time.Date(2024, time.March, 17, 23, 0, 0, 0, loc))
+ if truncated {
+ t.Fatal("expansion was truncated")
+ }
+ for _, o := range got {
+ if h, m, _ := o.In(loc).Clock(); h != 9 || m != 0 {
+ t.Errorf("occurrence %s drifted off the 09:00 wall clock", o.In(loc))
+ }
+ }
+ if len(got) != 3 {
+ t.Errorf("want 3 occurrences, got %d", len(got))
+ }
+}
diff --git a/index/digest.go b/index/digest.go
new file mode 100644
index 0000000..72be6fd
--- /dev/null
+++ b/index/digest.go
@@ -0,0 +1,76 @@
+package index
+
+import (
+ "sort"
+ "strings"
+ "time"
+
+ "github.com/dcadolph/midden/internal/util"
+)
+
+// monthLayout renders a calendar month as YYYY-MM.
+const monthLayout = "2006-01"
+
+// Digest is a whole-corpus summary computed from counts over every indexed
+// entry. Nearest-neighbor retrieval answers questions about a specific memory
+// but cannot answer questions about the record as a whole, so aggregate shape
+// is measured here and supplied alongside whatever entries were retrieved.
+type Digest struct {
+ // Entries is the number of indexed entries inside the range.
+ Entries int `json:"entries"`
+ // First is the earliest entry timestamp in range, zero when the range is empty.
+ First time.Time `json:"first,omitempty"`
+ // Last is the latest entry timestamp in range, zero when the range is empty.
+ Last time.Time `json:"last,omitempty"`
+ // TopTags is the tag histogram ordered by descending count then label.
+ TopTags []util.TagCount `json:"top_tags,omitempty"`
+ // Months is the per-month entry histogram ordered by ascending month.
+ Months []MonthCount `json:"months,omitempty"`
+}
+
+// MonthCount pairs a calendar month with the number of entries it holds.
+type MonthCount struct {
+ // Month is the calendar month in YYYY-MM form.
+ Month string `json:"month"`
+ // Count is the number of entries timestamped inside the month.
+ Count int `json:"count"`
+}
+
+// Digest summarizes every indexed entry inside the range. A zero from or to
+// leaves that end of the range open. TopTags is capped at topTags; a
+// non-positive value includes every tag.
+func (i *Index) Digest(from, to time.Time, topTags int) Digest {
+ var d Digest
+ tags := map[string]int{}
+ months := map[string]int{}
+ for _, e := range i.Entries {
+ if !inRange(e.Time, from, to) {
+ continue
+ }
+ d.Entries++
+ if d.First.IsZero() || e.Time.Before(d.First) {
+ d.First = e.Time
+ }
+ if e.Time.After(d.Last) {
+ d.Last = e.Time
+ }
+ for _, t := range e.Tags {
+ tags[strings.ToLower(t)]++
+ }
+ months[e.Time.Format(monthLayout)]++
+ }
+ d.TopTags = util.SortedCounts(tags, topTags)
+ d.Months = sortedMonths(months)
+ return d
+}
+
+// sortedMonths renders the month histogram in calendar order. The YYYY-MM
+// layout is zero-padded, so ascending lexical order is ascending calendar order.
+func sortedMonths(counts map[string]int) []MonthCount {
+ out := make([]MonthCount, 0, len(counts))
+ for month, n := range counts {
+ out = append(out, MonthCount{Month: month, Count: n})
+ }
+ sort.Slice(out, func(i, j int) bool { return out[i].Month < out[j].Month })
+ return out
+}
diff --git a/index/digest_test.go b/index/digest_test.go
new file mode 100644
index 0000000..eff8399
--- /dev/null
+++ b/index/digest_test.go
@@ -0,0 +1,134 @@
+package index
+
+import (
+ "fmt"
+ "strings"
+ "testing"
+ "time"
+
+ "github.com/google/go-cmp/cmp"
+ "github.com/google/go-cmp/cmp/cmpopts"
+
+ "github.com/dcadolph/midden/internal/util"
+)
+
+// day returns a local timestamp for the given calendar date at noon.
+func day(year int, month time.Month, d int) time.Time {
+ return time.Date(year, month, d, 12, 0, 0, 0, time.Local)
+}
+
+// digestFixture is a small corpus spanning three months across two years.
+func digestFixture() *Index {
+ return &Index{
+ Provider: "test",
+ Dim: 2,
+ Entries: []Entry{
+ {Time: day(2024, time.March, 3), Tags: []string{"work"}, Body: "alpha", Embedding: []float32{1, 0}},
+ {Time: day(2024, time.March, 20), Tags: []string{"Work", "travel"}, Body: "beta", Embedding: []float32{0, 1}},
+ {Time: day(2024, time.May, 9), Tags: []string{"home"}, Body: "gamma", Embedding: []float32{1, 1}},
+ {Time: day(2025, time.January, 15), Tags: []string{"work"}, Body: "delta", Embedding: []float32{-1, 0}},
+ },
+ }
+}
+
+func TestDigest(t *testing.T) {
+ t.Parallel()
+ tests := []struct {
+ From time.Time
+ To time.Time
+ WantSpan string
+ Want Digest
+ TopTags int
+ }{{ // Test 0: Open range covers the whole corpus and folds tag case.
+ TopTags: 0,
+ Want: Digest{
+ Entries: 4,
+ First: day(2024, time.March, 3),
+ Last: day(2025, time.January, 15),
+ TopTags: []util.TagCount{{Tag: "work", Count: 3}, {Tag: "home", Count: 1}, {Tag: "travel", Count: 1}},
+ Months: []MonthCount{
+ {Month: "2024-03", Count: 2},
+ {Month: "2024-05", Count: 1},
+ {Month: "2025-01", Count: 1},
+ },
+ },
+ }, { // Test 1: A bounded range counts only the entries inside it.
+ From: day(2024, time.March, 1),
+ To: day(2024, time.March, 31),
+ TopTags: 0,
+ Want: Digest{
+ Entries: 2,
+ First: day(2024, time.March, 3),
+ Last: day(2024, time.March, 20),
+ TopTags: []util.TagCount{{Tag: "work", Count: 2}, {Tag: "travel", Count: 1}},
+ Months: []MonthCount{{Month: "2024-03", Count: 2}},
+ },
+ }, { // Test 2: An open start bounds only the far end.
+ To: day(2024, time.March, 31),
+ TopTags: 1,
+ Want: Digest{
+ Entries: 2,
+ First: day(2024, time.March, 3),
+ Last: day(2024, time.March, 20),
+ TopTags: []util.TagCount{{Tag: "work", Count: 2}},
+ Months: []MonthCount{{Month: "2024-03", Count: 2}},
+ },
+ }, { // Test 3: A range holding no entries reports zero rather than the whole corpus.
+ From: day(2030, time.January, 1),
+ TopTags: 0,
+ Want: Digest{},
+ }}
+ for testNum, test := range tests {
+ t.Run(fmt.Sprintf("test %d", testNum), func(t *testing.T) {
+ t.Parallel()
+ got := digestFixture().Digest(test.From, test.To, test.TopTags)
+ if diff := cmp.Diff(test.Want, got, cmpopts.EquateEmpty()); diff != "" {
+ t.Errorf("mismatch (-want +got):\n%s", diff)
+ }
+ })
+ }
+}
+
+func TestDigestEmptyIndex(t *testing.T) {
+ t.Parallel()
+ got := (&Index{}).Digest(time.Time{}, time.Time{}, 10)
+ if diff := cmp.Diff(Digest{}, got, cmpopts.EquateEmpty()); diff != "" {
+ t.Errorf("mismatch (-want +got):\n%s", diff)
+ }
+}
+
+func TestEmbeddingsByContent(t *testing.T) {
+ t.Parallel()
+ idx := &Index{Entries: []Entry{
+ {Time: day(2024, time.March, 3), Body: "alpha", Embedding: []float32{1, 0}},
+ {Time: day(2024, time.March, 4), Body: "beta", Embedding: []float32{0, 1}},
+ // Same text on a different day shares a vector, because only the body is
+ // ever embedded.
+ {Time: day(2025, time.March, 3), Body: "alpha", Embedding: []float32{1, 0}},
+ // An entry still awaiting its vector contributes nothing to reuse.
+ {Time: day(2025, time.March, 4), Body: "gamma"},
+ }}
+ got := idx.EmbeddingsByContent()
+ if len(got) != 2 {
+ t.Fatalf("want 2 reusable vectors, got %d", len(got))
+ }
+ if diff := cmp.Diff([]float32{1, 0}, got[ContentHash("alpha")]); diff != "" {
+ t.Errorf("mismatch (-want +got):\n%s", diff)
+ }
+ if _, ok := got[ContentHash("gamma")]; ok {
+ t.Error("want no entry for a body with no embedding")
+ }
+}
+
+func TestContentHashIsStableAndDistinct(t *testing.T) {
+ t.Parallel()
+ // Built rather than written twice, so this checks equal text rather than one
+ // expression compared against itself.
+ rebuilt := "alp" + strings.Repeat("h", 1) + "a"
+ if ContentHash("alpha") != ContentHash(rebuilt) {
+ t.Error("want a stable hash for equal text")
+ }
+ if ContentHash("alpha") == ContentHash("alphb") {
+ t.Error("want different hashes for different text")
+ }
+}
diff --git a/index/index.go b/index/index.go
index 76b2acf..70a0702 100644
--- a/index/index.go
+++ b/index/index.go
@@ -3,6 +3,8 @@ package index
import (
"bytes"
+ "crypto/sha256"
+ "encoding/hex"
"encoding/json"
"fmt"
"math"
@@ -22,7 +24,7 @@ type Entry struct {
// Body is the entry text. Stored verbatim so recall can quote it back.
Body string `json:"body"`
// Embedding is the vector representation of Body.
- Embedding []float32 `json:"embedding"`
+ Embedding Vector `json:"embedding"`
}
// Index is the persistent vector store kept inside the vault root.
@@ -70,16 +72,55 @@ func (i *Index) Encode() ([]byte, error) {
return buf.Bytes(), nil
}
+// ContentHash returns the key that matches an entry body to a cached
+// embedding. Only the body is ever embedded, so two entries with the same text
+// can share a vector no matter how their timestamps or tags differ.
+func ContentHash(body string) string {
+ sum := sha256.Sum256([]byte(body))
+ return hex.EncodeToString(sum[:])
+}
+
+// EmbeddingsByContent maps each indexed body's content hash to its embedding so
+// a rebuild can reuse the vectors for text that has not changed. Rebuilding a
+// vault holding years of backfilled history costs one provider call per batch
+// of new entries this way, rather than re-embedding the whole corpus every time
+// a single day is added.
+func (i *Index) EmbeddingsByContent() map[string][]float32 {
+ out := make(map[string][]float32, len(i.Entries))
+ for _, e := range i.Entries {
+ if len(e.Embedding) == 0 {
+ continue
+ }
+ out[ContentHash(e.Body)] = e.Embedding
+ }
+ return out
+}
+
// Search returns the top-k entries by cosine similarity to the query vector.
// A non-positive k returns every entry ranked.
func (i *Index) Search(query []float32, k int) []Match {
+ return i.SearchRange(query, k, time.Time{}, time.Time{})
+}
+
+// SearchRange returns the top-k entries by cosine similarity to the query
+// vector, considering only entries timestamped inside the range. A zero from or
+// to leaves that end of the range open, and a non-positive k returns every
+// candidate ranked. Scoping by time before ranking keeps a question about one
+// period from matching a semantically similar entry years away from it.
+func (i *Index) SearchRange(query []float32, k int, from, to time.Time) []Match {
if len(i.Entries) == 0 {
return nil
}
matches := make([]Match, 0, len(i.Entries))
for _, e := range i.Entries {
+ if !inRange(e.Time, from, to) {
+ continue
+ }
matches = append(matches, Match{Entry: e, Score: cosine(query, e.Embedding)})
}
+ if len(matches) == 0 {
+ return nil
+ }
sort.Slice(matches, func(a, b int) bool { return matches[a].Score > matches[b].Score })
if k > 0 && len(matches) > k {
matches = matches[:k]
@@ -87,6 +128,33 @@ func (i *Index) Search(query []float32, k int) []Match {
return matches
}
+// InRange returns every entry timestamped inside the range ordered by ascending
+// timestamp. A zero from or to leaves that end of the range open. Questions
+// about a period are answered from the whole period rather than from its
+// nearest neighbors, so callers take the full slice instead of a ranked head.
+func (i *Index) InRange(from, to time.Time) []Entry {
+ out := make([]Entry, 0, len(i.Entries))
+ for _, e := range i.Entries {
+ if inRange(e.Time, from, to) {
+ out = append(out, e)
+ }
+ }
+ sort.Slice(out, func(a, b int) bool { return out[a].Time.Before(out[b].Time) })
+ return out
+}
+
+// inRange reports whether t falls inside the closed range, treating a zero
+// bound as open.
+func inRange(t, from, to time.Time) bool {
+ if !from.IsZero() && t.Before(from) {
+ return false
+ }
+ if !to.IsZero() && t.After(to) {
+ return false
+ }
+ return true
+}
+
// cosine returns the cosine similarity between two equal-length vectors.
// Vectors of different length or zero magnitude return zero.
func cosine(a, b []float32) float32 {
diff --git a/index/scope_test.go b/index/scope_test.go
new file mode 100644
index 0000000..8060a97
--- /dev/null
+++ b/index/scope_test.go
@@ -0,0 +1,106 @@
+package index
+
+import (
+ "fmt"
+ "testing"
+ "time"
+
+ "github.com/google/go-cmp/cmp"
+ "github.com/google/go-cmp/cmp/cmpopts"
+)
+
+func TestSearchRange(t *testing.T) {
+ t.Parallel()
+ tests := []struct {
+ From time.Time
+ To time.Time
+ WantBodies []string
+ K int
+ }{{ // Test 0: An open range behaves like an unscoped search.
+ K: 4,
+ WantBodies: []string{"alpha", "gamma", "beta", "delta"},
+ }, { // Test 1: Scoping excludes a closer match that falls outside the range.
+ From: day(2024, time.May, 1),
+ K: 4,
+ WantBodies: []string{"gamma", "delta"},
+ }, { // Test 2: A closed range keeps only entries inside both bounds.
+ From: day(2024, time.March, 1),
+ To: day(2024, time.March, 31),
+ K: 4,
+ WantBodies: []string{"alpha", "beta"},
+ }, { // Test 3: k truncates the ranked head after scoping.
+ K: 2,
+ WantBodies: []string{"alpha", "gamma"},
+ }, { // Test 4: A non-positive k returns every candidate in range.
+ From: day(2024, time.March, 1),
+ To: day(2024, time.March, 31),
+ WantBodies: []string{"alpha", "beta"},
+ }, { // Test 5: A range holding no entries returns nothing.
+ From: day(2030, time.January, 1),
+ K: 4,
+ WantBodies: nil,
+ }}
+ for testNum, test := range tests {
+ t.Run(fmt.Sprintf("test %d", testNum), func(t *testing.T) {
+ t.Parallel()
+ matches := digestFixture().SearchRange([]float32{1, 0}, test.K, test.From, test.To)
+ got := make([]string, len(matches))
+ for i, m := range matches {
+ got[i] = m.Entry.Body
+ }
+ if diff := cmp.Diff(test.WantBodies, got, cmpopts.EquateEmpty()); diff != "" {
+ t.Errorf("mismatch (-want +got):\n%s", diff)
+ }
+ })
+ }
+}
+
+func TestInRange(t *testing.T) {
+ t.Parallel()
+ tests := []struct {
+ From time.Time
+ To time.Time
+ WantBodies []string
+ }{{ // Test 0: An open range returns the whole corpus chronologically.
+ WantBodies: []string{"alpha", "beta", "gamma", "delta"},
+ }, { // Test 1: Bounds are inclusive at both ends.
+ From: day(2024, time.March, 3),
+ To: day(2024, time.May, 9),
+ WantBodies: []string{"alpha", "beta", "gamma"},
+ }, { // Test 2: An open end bounds only the near side.
+ From: day(2025, time.January, 1),
+ WantBodies: []string{"delta"},
+ }, { // Test 3: A range holding no entries returns nothing.
+ To: day(2000, time.January, 1),
+ WantBodies: nil,
+ }}
+ for testNum, test := range tests {
+ t.Run(fmt.Sprintf("test %d", testNum), func(t *testing.T) {
+ t.Parallel()
+ entries := digestFixture().InRange(test.From, test.To)
+ got := make([]string, len(entries))
+ for i, e := range entries {
+ got[i] = e.Body
+ }
+ if diff := cmp.Diff(test.WantBodies, got, cmpopts.EquateEmpty()); diff != "" {
+ t.Errorf("mismatch (-want +got):\n%s", diff)
+ }
+ })
+ }
+}
+
+func TestInRangeOrdersUnsortedEntries(t *testing.T) {
+ t.Parallel()
+ idx := &Index{Entries: []Entry{
+ {Time: day(2025, time.June, 1), Body: "late"},
+ {Time: day(2020, time.June, 1), Body: "early"},
+ {Time: day(2023, time.June, 1), Body: "middle"},
+ }}
+ got := make([]string, 0, 3)
+ for _, e := range idx.InRange(time.Time{}, time.Time{}) {
+ got = append(got, e.Body)
+ }
+ if diff := cmp.Diff([]string{"early", "middle", "late"}, got); diff != "" {
+ t.Errorf("mismatch (-want +got):\n%s", diff)
+ }
+}
diff --git a/index/vector.go b/index/vector.go
new file mode 100644
index 0000000..d7ada20
--- /dev/null
+++ b/index/vector.go
@@ -0,0 +1,60 @@
+package index
+
+import (
+ "encoding/base64"
+ "encoding/binary"
+ "encoding/json"
+ "fmt"
+ "math"
+)
+
+// float32Bytes is the width of one embedding component on the wire.
+const float32Bytes = 4
+
+// Vector is an embedding, stored on disk as a base64 blob of little-endian
+// float32 values rather than as a JSON array of numbers.
+//
+// The encoding matters at the scale a backfilled vault reaches. Measured over
+// ten thousand entries at 1536 dimensions, the index is 185 MB written as JSON
+// number arrays and 80 MB written as blobs. Decoding is the larger win, because
+// the whole index is read on every recall and chat: 2.8 seconds of float
+// parsing becomes 0.75 seconds of copying.
+type Vector []float32
+
+// MarshalJSON renders the vector as a base64 blob. A nil vector marshals as
+// JSON null so an entry still awaiting its embedding round trips unchanged.
+func (v Vector) MarshalJSON() ([]byte, error) {
+ if v == nil {
+ return []byte("null"), nil
+ }
+ raw := make([]byte, len(v)*float32Bytes)
+ for i, f := range v {
+ binary.LittleEndian.PutUint32(raw[i*float32Bytes:], math.Float32bits(f))
+ }
+ return json.Marshal(base64.StdEncoding.EncodeToString(raw))
+}
+
+// UnmarshalJSON reads a vector written by MarshalJSON.
+func (v *Vector) UnmarshalJSON(data []byte) error {
+ if string(data) == "null" {
+ *v = nil
+ return nil
+ }
+ var encoded string
+ if err := json.Unmarshal(data, &encoded); err != nil {
+ return fmt.Errorf("decode embedding: expected a base64 blob: run `midden reindex` to rebuild the index: %w", err)
+ }
+ raw, err := base64.StdEncoding.DecodeString(encoded)
+ if err != nil {
+ return fmt.Errorf("decode embedding: %w", err)
+ }
+ if len(raw)%float32Bytes != 0 {
+ return fmt.Errorf("decode embedding: %d bytes is not a whole number of float32 values", len(raw))
+ }
+ out := make(Vector, len(raw)/float32Bytes)
+ for i := range out {
+ out[i] = math.Float32frombits(binary.LittleEndian.Uint32(raw[i*float32Bytes:]))
+ }
+ *v = out
+ return nil
+}
diff --git a/index/vector_test.go b/index/vector_test.go
new file mode 100644
index 0000000..22d9166
--- /dev/null
+++ b/index/vector_test.go
@@ -0,0 +1,102 @@
+package index
+
+import (
+ "encoding/json"
+ "math"
+ "testing"
+
+ "github.com/google/go-cmp/cmp"
+)
+
+func TestVectorRoundTrip(t *testing.T) {
+ t.Parallel()
+ tests := []struct {
+ Name string
+ In Vector
+ }{
+ {Name: "typical", In: Vector{0.1, -0.25, 1, -1, 0}},
+ {Name: "empty", In: Vector{}},
+ {Name: "extremes", In: Vector{math.MaxFloat32, -math.MaxFloat32, math.SmallestNonzeroFloat32}},
+ }
+ for _, test := range tests {
+ t.Run(test.Name, func(t *testing.T) {
+ t.Parallel()
+ data, err := json.Marshal(test.In)
+ if err != nil {
+ t.Fatalf("Marshal: %v", err)
+ }
+ var got Vector
+ if err := json.Unmarshal(data, &got); err != nil {
+ t.Fatalf("Unmarshal: %v", err)
+ }
+ // The blob is exact rather than approximate: the bits go out and come
+ // back unchanged, so search scores do not drift with the file format.
+ if diff := cmp.Diff(test.In, got); diff != "" {
+ t.Errorf("mismatch (-want +got):\n%s", diff)
+ }
+ })
+ }
+}
+
+func TestVectorNilRoundTrip(t *testing.T) {
+ t.Parallel()
+ data, err := json.Marshal(Vector(nil))
+ if err != nil {
+ t.Fatalf("Marshal: %v", err)
+ }
+ if string(data) != "null" {
+ t.Errorf("want null for an entry with no embedding yet, got %s", data)
+ }
+ var got Vector
+ if err := json.Unmarshal(data, &got); err != nil {
+ t.Fatalf("Unmarshal: %v", err)
+ }
+ if got != nil {
+ t.Errorf("want nil, got %v", got)
+ }
+}
+
+func TestVectorRejectsBadInput(t *testing.T) {
+ t.Parallel()
+ tests := []struct {
+ Name string
+ In string
+ }{
+ {Name: "json array from an older index", In: "[0.1,0.2,0.3]"},
+ {Name: "not base64", In: `"not base64!!"`},
+ {Name: "truncated float", In: `"AAAAAAA="`},
+ }
+ for _, test := range tests {
+ t.Run(test.Name, func(t *testing.T) {
+ t.Parallel()
+ var got Vector
+ if err := json.Unmarshal([]byte(test.In), &got); err == nil {
+ t.Errorf("want an error for %s, got %v", test.Name, got)
+ }
+ })
+ }
+}
+
+func TestIndexRoundTripsThroughBlobVectors(t *testing.T) {
+ t.Parallel()
+ idx := &Index{
+ Provider: "test",
+ Dim: 3,
+ Entries: []Entry{
+ {Body: "alpha", Embedding: Vector{1, 0, 0}},
+ {Body: "beta", Embedding: Vector{0, 0.5, -0.5}},
+ {Body: "not embedded yet"},
+ },
+ }
+ data, err := idx.Encode()
+ if err != nil {
+ t.Fatalf("Encode: %v", err)
+ }
+ got, err := Decode(data)
+ if err != nil {
+ t.Fatalf("Decode: %v", err)
+ }
+ if diff := cmp.Diff(idx.Entries, got.Entries); diff != "" {
+ t.Errorf("mismatch (-want +got):\n%s", diff)
+ }
+}
diff --git a/internal/interview/question.go b/internal/interview/question.go
new file mode 100644
index 0000000..53155b7
--- /dev/null
+++ b/internal/interview/question.go
@@ -0,0 +1,273 @@
+// Package interview turns gaps in a record into questions worth answering.
+//
+// Backfill can reconstruct where a person was and what they produced, because
+// calendars and commit logs already exist. It cannot reconstruct what they
+// thought, since nothing recorded that at the time. The only way past that
+// ceiling is to ask, and the only questions worth asking are the ones the
+// record can prove are missing: a commitment held for years that stopped
+// without explanation, a thread that appeared from nowhere, a day where two
+// parts of a life collided.
+//
+// Every question here is derived from arithmetic over the record rather than
+// generated by a model, so a question is never asked about something that did
+// not happen.
+package interview
+
+import (
+ "crypto/sha256"
+ "encoding/hex"
+ "fmt"
+ "sort"
+ "time"
+
+ "github.com/dcadolph/midden/internal/weave"
+)
+
+// Kind is the sort of gap a question addresses.
+type Kind string
+
+// The kinds of gap worth asking about.
+const (
+ // KindEnded asks why a long-running thread stopped.
+ KindEnded Kind = "ended"
+ // KindBegan asks how a new thread came about.
+ KindBegan Kind = "began"
+ // KindCrossing asks about a day where separate parts of a life met.
+ KindCrossing Kind = "crossing"
+ // KindGap asks what was happening while the record itself went quiet.
+ KindGap Kind = "gap"
+)
+
+// Question is one thing the record cannot answer about itself.
+type Question struct {
+ // ID identifies the question so it is never asked twice.
+ ID string
+ // Kind is the sort of gap it addresses.
+ Kind Kind
+ // Prompt is the question as it is put to the person.
+ Prompt string
+ // Context is the evidence the question rests on, shown alongside it.
+ Context string
+ // When is the date the question concerns, used to file the answer.
+ When time.Time
+ // Weight orders questions, heaviest first.
+ Weight int
+}
+
+// Options tune question generation.
+type Options struct {
+ // Now is the observation date.
+ Now time.Time
+ // MinWeight is the smallest thread weight worth asking about, which keeps
+ // incidental patterns from producing questions.
+ MinWeight int
+ // MinSpanDays is how long a thread must have run before its ending is worth
+ // explaining. A school-year reminder repeated for a term is logistics, not a
+ // commitment, and asking why it stopped mistakes a to-do list for a life.
+ MinSpanDays int
+}
+
+// DefaultOptions returns generation settings suited to a personal record.
+func DefaultOptions(now time.Time) Options {
+ return Options{Now: now, MinWeight: 300, MinSpanDays: 180}
+}
+
+// Generate builds the questions a record supports, heaviest first, skipping any
+// whose ID is already answered.
+func Generate(
+ threads []weave.Thread,
+ crossings []weave.Overlap,
+ gaps []weave.Gap,
+ answered map[string]bool,
+ opts Options,
+) []Question {
+ var out []Question
+ add := func(q Question) {
+ if q.ID == "" || answered[q.ID] {
+ return
+ }
+ out = append(out, q)
+ }
+ for _, t := range threads {
+ if t.Weight() < opts.MinWeight {
+ continue
+ }
+ switch t.Status {
+ case weave.Ended:
+ if t.SpanDays >= opts.MinSpanDays {
+ add(endedQuestion(t))
+ }
+ case weave.Emerging:
+ add(beganQuestion(t))
+ case weave.Dormant:
+ // Between seasons, not over. Nothing to explain.
+ case weave.Ongoing:
+ // A thread still running is not a gap. Asking about one produces
+ // questions about groceries and school holidays, because frequency
+ // alone cannot tell a commitment that mattered from a chore that
+ // recurred, and a vapid question teaches the person to ignore the
+ // prompt entirely.
+ }
+ }
+ for _, c := range crossings {
+ add(crossingQuestion(c))
+ }
+ // One silence per source at most, the longest. Three questions about what is
+ // essentially one intermittent quiet stretch is over-asking, and over-asking
+ // is how a prompt earns being ignored.
+ seenSource := map[string]bool{}
+ for _, g := range gaps {
+ if seenSource[g.Source] {
+ continue
+ }
+ seenSource[g.Source] = true
+ add(gapQuestion(g))
+ }
+ // Silences are asked first regardless of weight. A gap and a thread are not
+ // measured on the same scale, and no arithmetic makes them comparable: one
+ // counts occurrences of a commitment, the other counts months of nothing.
+ // A stretch of years the record cannot account for is the larger hole, so
+ // kind decides the order and weight only breaks ties inside a kind.
+ sort.Slice(out, func(i, j int) bool {
+ pi, pj := priority(out[i].Kind), priority(out[j].Kind)
+ if pi != pj {
+ return pi < pj
+ }
+ if out[i].Weight != out[j].Weight {
+ return out[i].Weight > out[j].Weight
+ }
+ return out[i].ID < out[j].ID
+ })
+ return out
+}
+
+// priority orders the kinds of gap, lowest first.
+func priority(k Kind) int {
+ switch k {
+ case KindGap:
+ return 0
+ case KindEnded:
+ return 1
+ case KindCrossing:
+ return 2
+ case KindBegan:
+ return 3
+ }
+ return 4
+}
+
+// endedQuestion asks about a thread that stopped. An ending is the most
+// valuable thing to ask about, because it is the one event a person never
+// records: there is a last time, and nobody knows it is the last time.
+func endedQuestion(t weave.Thread) Question {
+ return Question{
+ ID: id(KindEnded, t.Key),
+ Kind: KindEnded,
+ Prompt: fmt.Sprintf("%s stopped. What happened?", t.Label),
+ Context: fmt.Sprintf("%d times over %s, ending %s. Nothing since, %s ago.",
+ t.Count, approxYears(t.SpanDays), t.Last.Format(dateLayout), approxMonths(t.SilentDays)),
+ When: t.Last,
+ Weight: t.Weight(),
+ }
+}
+
+// beganQuestion asks how something new came about.
+func beganQuestion(t weave.Thread) Question {
+ return Question{
+ ID: id(KindBegan, t.Key),
+ Kind: KindBegan,
+ Prompt: fmt.Sprintf("%s is new. How did that come about?", t.Label),
+ Context: fmt.Sprintf("%d times since %s.", t.Count, t.First.Format(dateLayout)),
+ When: t.First,
+ Weight: t.Weight(),
+ }
+}
+
+// crossingQuestion asks about a day where separate sources both recorded
+// something, which is the only place the record can point at a tension it
+// cannot explain.
+func crossingQuestion(c weave.Overlap) Question {
+ day := c.Day.Format(dateLayout)
+ total := 0
+ for _, n := range c.Counts {
+ total += n
+ }
+ var lines string
+ for _, src := range sortedKeys(c.Headlines) {
+ lines += fmt.Sprintf("%s: %s (%d)\n", src, c.Headlines[src], c.Counts[src])
+ }
+ return Question{
+ ID: id(KindCrossing, day),
+ Kind: KindCrossing,
+ Prompt: fmt.Sprintf("On %s two parts of your life ran at once. What was going on?", day),
+ Context: lines,
+ When: c.Day,
+ Weight: total,
+ }
+}
+
+// gapQuestion asks about a stretch where the record fell silent. A silence is
+// the largest thing a record can be missing and the least visible from inside
+// it: a person notices a class ending, never that years went unrecorded.
+func gapQuestion(g weave.Gap) Question {
+ span := fmt.Sprintf("%s to %s", g.From.Format(monthLabel), g.To.Format(monthLabel))
+ subject := "Your record"
+ if g.Source != "" {
+ subject = "Your " + g.Source
+ }
+ return Question{
+ ID: id(KindGap, g.Source+"|"+span),
+ Kind: KindGap,
+ Prompt: fmt.Sprintf("%s goes quiet for %s, from %s. What was happening then?",
+ subject, approxMonths(g.Months*30), span),
+ Context: fmt.Sprintf("%d entries across %d months, against about %.0f a month before and %.0f after.",
+ g.Entries, g.Months, g.Before, g.After),
+ When: g.From,
+ Weight: g.Months,
+ }
+}
+
+// monthLabel renders a month as YYYY-MM.
+const monthLabel = "2006-01"
+
+// dateLayout renders a date as YYYY-MM-DD.
+const dateLayout = "2006-01-02"
+
+// id derives a stable short identifier for a question, so the same gap always
+// produces the same key and an answered question is never asked again.
+func id(kind Kind, subject string) string {
+ if subject == "" {
+ return ""
+ }
+ sum := sha256.Sum256([]byte(string(kind) + "\x00" + subject))
+ return string(kind) + "-" + hex.EncodeToString(sum[:4])
+}
+
+// sortedKeys returns map keys in a stable order.
+func sortedKeys(m map[string]string) []string {
+ out := make([]string, 0, len(m))
+ for k := range m {
+ out = append(out, k)
+ }
+ sort.Strings(out)
+ return out
+}
+
+// approxYears renders a day count as an approximate number of years or months.
+func approxYears(days int) string {
+ if days < 365 {
+ return approxMonths(days)
+ }
+ return fmt.Sprintf("%.1f years", float64(days)/365)
+}
+
+// approxMonths renders a day count as an approximate number of months or days.
+func approxMonths(days int) string {
+ switch {
+ case days < 45:
+ return fmt.Sprintf("%d days", days)
+ case days < 365:
+ return fmt.Sprintf("%d months", days/30)
+ }
+ return fmt.Sprintf("%.1f years", float64(days)/365)
+}
diff --git a/internal/interview/question_test.go b/internal/interview/question_test.go
new file mode 100644
index 0000000..eaf4137
--- /dev/null
+++ b/internal/interview/question_test.go
@@ -0,0 +1,192 @@
+package interview
+
+import (
+ "strings"
+ "testing"
+ "time"
+
+ "github.com/dcadolph/midden/internal/weave"
+)
+
+// observed is a fixed observation date so generation never depends on the clock.
+var observed = time.Date(2026, time.August, 26, 12, 0, 0, 0, time.Local)
+
+// day returns a local date.
+func day(year int, month time.Month, d int) time.Time {
+ return time.Date(year, month, d, 12, 0, 0, 0, time.Local)
+}
+
+// endedThread builds a heavy thread that has gone quiet.
+func endedThread(key, label string) weave.Thread {
+ return weave.Thread{
+ Key: key, Label: label, Count: 216,
+ First: day(2023, time.January, 10), Last: day(2025, time.August, 7),
+ MedianGap: 4, SilentDays: 384, SpanDays: 940, Status: weave.Ended,
+ }
+}
+
+func TestGenerateAsksAboutEndings(t *testing.T) {
+ t.Parallel()
+ got := Generate([]weave.Thread{endedThread("arts martial william", "William- Martial arts")},
+ nil, nil, nil, DefaultOptions(observed))
+ if len(got) != 1 {
+ t.Fatalf("want one question, got %d", len(got))
+ }
+ q := got[0]
+ if q.Kind != KindEnded {
+ t.Errorf("want an ending question, got %s", q.Kind)
+ }
+ // The evidence has to travel with the question, or the person is being asked
+ // to remember the very thing they cannot.
+ for _, want := range []string{"216", "2025-08-07"} {
+ if !strings.Contains(q.Context, want) {
+ t.Errorf("want %q in the evidence, got %q", want, q.Context)
+ }
+ }
+}
+
+func TestGenerateSkipsOngoingThreads(t *testing.T) {
+ t.Parallel()
+ // A grocery run recurring for years is not a gap, and asking about it
+ // teaches the person to ignore the prompt.
+ chore := weave.Thread{
+ Key: "curbside heb", Label: "H-E-B curbside", Count: 239,
+ First: day(2022, time.October, 1), Last: observed,
+ MedianGap: 4, SpanDays: 1400, Status: weave.Ongoing,
+ }
+ if got := Generate([]weave.Thread{chore}, nil, nil, nil, DefaultOptions(observed)); len(got) != 0 {
+ t.Errorf("want nothing asked about a live thread, got %q", got[0].Prompt)
+ }
+}
+
+func TestGenerateSkipsLightThreads(t *testing.T) {
+ t.Parallel()
+ light := weave.Thread{
+ Key: "dentist", Label: "Dentist", Count: 3,
+ First: day(2024, time.March, 1), Last: day(2024, time.May, 1),
+ MedianGap: 30, SilentDays: 800, SpanDays: 61, Status: weave.Ended,
+ }
+ if got := Generate([]weave.Thread{light}, nil, nil, nil, DefaultOptions(observed)); len(got) != 0 {
+ t.Errorf("want incidental patterns ignored, got %q", got[0].Prompt)
+ }
+}
+
+func TestGenerateNeverRepeatsAnAnsweredQuestion(t *testing.T) {
+ t.Parallel()
+ th := endedThread("arts martial william", "William- Martial arts")
+ first := Generate([]weave.Thread{th}, nil, nil, nil, DefaultOptions(observed))
+ if len(first) != 1 {
+ t.Fatalf("want one question, got %d", len(first))
+ }
+ answered := map[string]bool{first[0].ID: true}
+ // Being asked again about something already answered is the fastest way to
+ // make the prompt worthless.
+ if again := Generate([]weave.Thread{th}, nil, nil, answered, DefaultOptions(observed)); len(again) != 0 {
+ t.Errorf("want an answered question retired, got %q", again[0].Prompt)
+ }
+}
+
+func TestQuestionIDIsStableAcrossRuns(t *testing.T) {
+ t.Parallel()
+ th := endedThread("arts martial william", "William- Martial arts")
+ a := Generate([]weave.Thread{th}, nil, nil, nil, DefaultOptions(observed))[0].ID
+ b := Generate([]weave.Thread{th}, nil, nil, nil, DefaultOptions(observed))[0].ID
+ if a != b {
+ t.Errorf("want a stable id so answers keep matching, got %q then %q", a, b)
+ }
+ other := endedThread("hip hannah hop", "Hannah- hip hop")
+ if c := Generate([]weave.Thread{other}, nil, nil, nil, DefaultOptions(observed))[0].ID; c == a {
+ t.Error("want different gaps to carry different ids")
+ }
+}
+
+func TestGenerateOrdersHeaviestFirst(t *testing.T) {
+ t.Parallel()
+ heavy := endedThread("arts martial william", "William- Martial arts")
+ lighter := weave.Thread{
+ Key: "acro hannah", Label: "Hannah- Acro", Count: 42,
+ First: day(2023, time.March, 1), Last: day(2024, time.April, 12),
+ MedianGap: 7, SilentDays: 866, SpanDays: 408, Status: weave.Ended,
+ }
+ got := Generate([]weave.Thread{lighter, heavy}, nil, nil, nil, DefaultOptions(observed))
+ if len(got) != 2 {
+ t.Fatalf("want two questions, got %d", len(got))
+ }
+ if got[0].Prompt != heavy.Label+" stopped. What happened?" {
+ t.Errorf("want the bigger loss asked first, got %q", got[0].Prompt)
+ }
+}
+
+func TestCrossingQuestionCarriesBothSides(t *testing.T) {
+ t.Parallel()
+ c := weave.Overlap{
+ Day: day(2026, time.July, 30),
+ Counts: map[string]int{"calendar": 1, "git": 113},
+ Headlines: map[string]string{"calendar": "Dad- off work/ vacation", "git": "Trim projects"},
+ }
+ got := Generate(nil, []weave.Overlap{c}, nil, nil, DefaultOptions(observed))
+ if len(got) != 1 {
+ t.Fatalf("want one crossing question, got %d", len(got))
+ }
+ // A crossing is only meaningful if both sides are shown; either alone is
+ // unremarkable.
+ for _, want := range []string{"off work", "Trim projects"} {
+ if !strings.Contains(got[0].Context, want) {
+ t.Errorf("want %q in the evidence, got %q", want, got[0].Context)
+ }
+ }
+}
+
+func TestGenerateEmptyRecord(t *testing.T) {
+ t.Parallel()
+ if got := Generate(nil, nil, nil, nil, DefaultOptions(observed)); len(got) != 0 {
+ t.Errorf("want nothing asked of an empty record, got %d", len(got))
+ }
+}
+
+func TestGenerateSkipsShortLivedThreads(t *testing.T) {
+ t.Parallel()
+ // A school-year reminder repeated for one term. It carries enough weight to
+ // pass the volume bar, but it was logistics, never a commitment, so asking
+ // why it stopped mistakes a to-do list for a life.
+ reminder := weave.Thread{
+ Key: "book class folder red return shirt wear william", Label: "William- wear red class shirt",
+ Count: 16, First: day(2024, time.January, 26), Last: day(2024, time.May, 10),
+ MedianGap: 7, SilentDays: 838, SpanDays: 105, Status: weave.Ended,
+ }
+ if got := Generate([]weave.Thread{reminder}, nil, nil, nil, DefaultOptions(observed)); len(got) != 0 {
+ t.Errorf("want a short reminder run ignored, got %q", got[0].Prompt)
+ }
+}
+
+func TestGenerateSkipsDormantThreads(t *testing.T) {
+ t.Parallel()
+ seasonal := weave.Thread{
+ Key: "haylie show spring", Label: "Haylie- spring show",
+ Count: 6, First: day(2023, time.April, 25), Last: day(2026, time.April, 25),
+ MedianGap: 1, MaxGap: 760, SilentDays: 123, SpanDays: 1096, Status: weave.Dormant,
+ }
+ // Asking a person to explain the end of something that has not ended asserts
+ // a false fact about their life.
+ if got := Generate([]weave.Thread{seasonal}, nil, nil, nil, DefaultOptions(observed)); len(got) != 0 {
+ t.Errorf("want a dormant thread left alone, got %q", got[0].Prompt)
+ }
+}
+
+func TestGapQuestionOutranksEveryThread(t *testing.T) {
+ t.Parallel()
+ heavy := endedThread("arts martial william", "William- Martial arts")
+ gap := weave.Gap{
+ From: day(2017, time.July, 1), To: day(2022, time.July, 31),
+ Months: 61, Entries: 6, Before: 4, After: 98,
+ }
+ got := Generate([]weave.Thread{heavy}, nil, []weave.Gap{gap}, nil, DefaultOptions(observed))
+ if len(got) != 2 {
+ t.Fatalf("want both questions, got %d", len(got))
+ }
+ // Years the record cannot account for is a larger hole than any one
+ // commitment ending, and the two are not measured on comparable scales.
+ if got[0].Kind != KindGap {
+ t.Errorf("want the silence asked first, got %s", got[0].Kind)
+ }
+}
diff --git a/internal/report/report.go b/internal/report/report.go
index 72ce53c..55798cc 100644
--- a/internal/report/report.go
+++ b/internal/report/report.go
@@ -9,6 +9,7 @@ import (
"strings"
"time"
+ "github.com/dcadolph/midden/internal/util"
"github.com/dcadolph/midden/internal/vault"
)
@@ -75,11 +76,11 @@ func Render(w io.Writer, data Data) error {
// Build prepares Data for Render by walking the vault.
// The heatmap covers the full year ending on the supplied reference date.
func Build(title string, v *vault.Vault, now time.Time, topTags int) (Data, error) {
- stats, err := v.ComputeStats(topTags)
+ stats, err := v.ComputeStats(topTags, now)
if err != nil {
return Data{}, fmt.Errorf("compute stats: %w", err)
}
- streak, err := v.Streak(now)
+ streak, err := v.Streak(now, vault.Entry.Authored)
if err != nil {
return Data{}, fmt.Errorf("compute streak: %w", err)
}
@@ -169,7 +170,7 @@ func buildDayRows(days []time.Time, counts map[string]dayMetric) []DayRow {
// buildTagBars scales tag counts to percent widths against the top tag.
// The top tag gets width 100 and every other tag gets at least width 4.
-func buildTagBars(tags []vault.TagCount) []TagBar {
+func buildTagBars(tags []util.TagCount) []TagBar {
if len(tags) == 0 {
return nil
}
diff --git a/internal/report/report_test.go b/internal/report/report_test.go
index 6c7e23a..b6542e0 100644
--- a/internal/report/report_test.go
+++ b/internal/report/report_test.go
@@ -10,6 +10,7 @@ import (
"github.com/google/go-cmp/cmp"
"github.com/google/go-cmp/cmp/cmpopts"
+ "github.com/dcadolph/midden/internal/util"
"github.com/dcadolph/midden/internal/vault"
)
@@ -189,23 +190,23 @@ func TestBuildDayRows(t *testing.T) {
func TestBuildTagBars(t *testing.T) {
t.Parallel()
tests := []struct {
- In []vault.TagCount
+ In []util.TagCount
WantBars []TagBar
}{{ // Test 0: The top tag gets width 100 and tiny tags get the minimum width 4.
- In: []vault.TagCount{{Tag: "work", Count: 50}, {Tag: "home", Count: 25}, {Tag: "gym", Count: 1}},
+ In: []util.TagCount{{Tag: "work", Count: 50}, {Tag: "home", Count: 25}, {Tag: "gym", Count: 1}},
WantBars: []TagBar{
{Tag: "work", Count: 50, Width: 100},
{Tag: "home", Count: 25, Width: 50},
{Tag: "gym", Count: 1, Width: 4},
},
}, { // Test 1: A single tag spans the full width.
- In: []vault.TagCount{{Tag: "solo", Count: 3}},
+ In: []util.TagCount{{Tag: "solo", Count: 3}},
WantBars: []TagBar{{Tag: "solo", Count: 3, Width: 100}},
}, { // Test 2: No tags yields no bars.
In: nil,
WantBars: nil,
}, { // Test 3: A zero top count yields zero widths.
- In: []vault.TagCount{{Tag: "ghost", Count: 0}},
+ In: []util.TagCount{{Tag: "ghost", Count: 0}},
WantBars: []TagBar{{Tag: "ghost", Count: 0, Width: 0}},
}}
for testNum, test := range tests {
diff --git a/internal/util/histogram.go b/internal/util/histogram.go
new file mode 100644
index 0000000..1464fc5
--- /dev/null
+++ b/internal/util/histogram.go
@@ -0,0 +1,30 @@
+package util
+
+import "sort"
+
+// TagCount pairs a label with the number of entries carrying it.
+type TagCount struct {
+ // Tag is the label without a leading hash character.
+ Tag string `json:"tag"`
+ // Count is the number of entries carrying the label.
+ Count int `json:"count"`
+}
+
+// SortedCounts renders a label histogram as a slice ordered by descending count
+// then ascending label. A non-positive limit returns every label.
+func SortedCounts(counts map[string]int, limit int) []TagCount {
+ out := make([]TagCount, 0, len(counts))
+ for tag, n := range counts {
+ out = append(out, TagCount{Tag: tag, Count: n})
+ }
+ sort.Slice(out, func(i, j int) bool {
+ if out[i].Count != out[j].Count {
+ return out[i].Count > out[j].Count
+ }
+ return out[i].Tag < out[j].Tag
+ })
+ if limit > 0 && len(out) > limit {
+ out = out[:limit]
+ }
+ return out
+}
diff --git a/internal/vault/append_all_test.go b/internal/vault/append_all_test.go
new file mode 100644
index 0000000..43e4288
--- /dev/null
+++ b/internal/vault/append_all_test.go
@@ -0,0 +1,147 @@
+package vault
+
+import (
+ "testing"
+ "time"
+
+ "github.com/google/go-cmp/cmp"
+)
+
+// batchEntry builds an entry on the given day at the given hour.
+func batchEntry(day, hour int, body string) Entry {
+ return Entry{Time: time.Date(2024, time.March, day, hour, 0, 0, 0, time.Local), Body: body}
+}
+
+// readBodies returns the entry bodies stored on the given day.
+func readBodies(t *testing.T, v *Vault, day time.Time) []string {
+ t.Helper()
+ entries, err := v.ReadDay(day)
+ if err != nil {
+ t.Fatalf("ReadDay: %v", err)
+ }
+ out := make([]string, len(entries))
+ for i, e := range entries {
+ out[i] = e.Body
+ }
+ return out
+}
+
+func TestAppendAllGroupsEntriesByDay(t *testing.T) {
+ t.Parallel()
+ v, err := Open(t.TempDir())
+ if err != nil {
+ t.Fatalf("Open: %v", err)
+ }
+ entries := []Entry{
+ batchEntry(4, 9, "monday first"),
+ batchEntry(5, 8, "tuesday"),
+ batchEntry(4, 17, "monday second"),
+ }
+ if err := v.AppendAll(entries); err != nil {
+ t.Fatalf("AppendAll: %v", err)
+ }
+ monday := time.Date(2024, time.March, 4, 0, 0, 0, 0, time.Local)
+ if diff := cmp.Diff([]string{"monday first", "monday second"}, readBodies(t, v, monday)); diff != "" {
+ t.Errorf("monday mismatch (-want +got):\n%s", diff)
+ }
+ tuesday := time.Date(2024, time.March, 5, 0, 0, 0, 0, time.Local)
+ if diff := cmp.Diff([]string{"tuesday"}, readBodies(t, v, tuesday)); diff != "" {
+ t.Errorf("tuesday mismatch (-want +got):\n%s", diff)
+ }
+}
+
+func TestAppendAllMatchesRepeatedAppend(t *testing.T) {
+ t.Parallel()
+ entries := []Entry{
+ batchEntry(4, 9, "one"),
+ batchEntry(4, 10, "two"),
+ batchEntry(6, 11, "three"),
+ }
+ batched, err := Open(t.TempDir())
+ if err != nil {
+ t.Fatalf("Open: %v", err)
+ }
+ if err := batched.AppendAll(entries); err != nil {
+ t.Fatalf("AppendAll: %v", err)
+ }
+ single, err := Open(t.TempDir())
+ if err != nil {
+ t.Fatalf("Open: %v", err)
+ }
+ for _, e := range entries {
+ if err := single.Append(e); err != nil {
+ t.Fatalf("Append: %v", err)
+ }
+ }
+ for _, day := range []int{4, 6} {
+ d := time.Date(2024, time.March, day, 0, 0, 0, 0, time.Local)
+ if diff := cmp.Diff(readBodies(t, single, d), readBodies(t, batched, d)); diff != "" {
+ t.Errorf("day %d differs from repeated Append (-single +batched):\n%s", day, diff)
+ }
+ }
+}
+
+func TestAppendAllOnEncryptedVault(t *testing.T) {
+ t.Parallel()
+ v, err := Open(t.TempDir())
+ if err != nil {
+ t.Fatalf("Open: %v", err)
+ }
+ if err := v.SetEncrypted(); err != nil {
+ t.Fatalf("SetEncrypted: %v", err)
+ }
+ v = v.WithPassphrase("correct horse battery staple")
+ if err := v.AppendAll([]Entry{batchEntry(4, 9, "sealed one"), batchEntry(4, 10, "sealed two")}); err != nil {
+ t.Fatalf("AppendAll: %v", err)
+ }
+ day := time.Date(2024, time.March, 4, 0, 0, 0, 0, time.Local)
+ if diff := cmp.Diff([]string{"sealed one", "sealed two"}, readBodies(t, v, day)); diff != "" {
+ t.Errorf("mismatch (-want +got):\n%s", diff)
+ }
+}
+
+func TestAppendAllAppendsToAnExistingDay(t *testing.T) {
+ t.Parallel()
+ v, err := Open(t.TempDir())
+ if err != nil {
+ t.Fatalf("Open: %v", err)
+ }
+ if err := v.Append(batchEntry(4, 8, "already here")); err != nil {
+ t.Fatalf("Append: %v", err)
+ }
+ if err := v.AppendAll([]Entry{batchEntry(4, 9, "added")}); err != nil {
+ t.Fatalf("AppendAll: %v", err)
+ }
+ day := time.Date(2024, time.March, 4, 0, 0, 0, 0, time.Local)
+ if diff := cmp.Diff([]string{"already here", "added"}, readBodies(t, v, day)); diff != "" {
+ t.Errorf("mismatch (-want +got):\n%s", diff)
+ }
+}
+
+func TestAppendAllRejectsEmptyBodyBeforeWriting(t *testing.T) {
+ t.Parallel()
+ v, err := Open(t.TempDir())
+ if err != nil {
+ t.Fatalf("Open: %v", err)
+ }
+ err = v.AppendAll([]Entry{batchEntry(4, 9, "good"), batchEntry(5, 9, " ")})
+ if err == nil {
+ t.Fatal("want an error for an empty body")
+ }
+ // Validation runs before any write so a bad batch never lands half-applied.
+ day := time.Date(2024, time.March, 4, 0, 0, 0, 0, time.Local)
+ if got := readBodies(t, v, day); len(got) != 0 {
+ t.Errorf("want nothing written, got %v", got)
+ }
+}
+
+func TestAppendAllEmptyBatchIsANoOp(t *testing.T) {
+ t.Parallel()
+ v, err := Open(t.TempDir())
+ if err != nil {
+ t.Fatalf("Open: %v", err)
+ }
+ if err := v.AppendAll(nil); err != nil {
+ t.Errorf("AppendAll(nil): %v", err)
+ }
+}
diff --git a/internal/vault/entry.go b/internal/vault/entry.go
index 519e9e0..455c1ca 100644
--- a/internal/vault/entry.go
+++ b/internal/vault/entry.go
@@ -3,6 +3,7 @@ package vault
import (
"fmt"
"os"
+ "sort"
"strings"
"time"
@@ -60,6 +61,95 @@ func (v *Vault) Append(entry Entry) error {
return nil
}
+// AppendAll writes every entry to its day file, grouping by day so each file is
+// touched once rather than once per entry. This matters for bulk ingestion: on
+// an encrypted vault Append decrypts and re-encrypts the whole day file per
+// call, so appending a backfill entry at a time costs one full crypt cycle per
+// event. Entries keep their given order within each day, every body is
+// validated before anything is written, and the batch runs under a single
+// advisory lock.
+func (v *Vault) AppendAll(entries []Entry) error {
+ if len(entries) == 0 {
+ return nil
+ }
+ for i, e := range entries {
+ if strings.TrimSpace(e.Body) == "" {
+ return fmt.Errorf("entry %d body is empty", i)
+ }
+ }
+ byDay := map[string][]Entry{}
+ var order []string
+ for _, e := range entries {
+ key := e.Time.Format("2006-01-02")
+ if _, ok := byDay[key]; !ok {
+ order = append(order, key)
+ }
+ byDay[key] = append(byDay[key], e)
+ }
+ sort.Strings(order)
+ lock, err := flock.Acquire(v.LockPath())
+ if err != nil {
+ return fmt.Errorf("acquire vault lock: %w", err)
+ }
+ defer func() { _ = lock.Close() }()
+ for _, key := range order {
+ if err := v.appendDay(byDay[key]); err != nil {
+ return fmt.Errorf("append %s: %w", key, err)
+ }
+ }
+ return nil
+}
+
+// appendDay writes a run of same-day entries to their day file. The caller
+// holds the vault lock.
+func (v *Vault) appendDay(entries []Entry) error {
+ path, err := v.EnsureDayFile(entries[0].Time)
+ if err != nil {
+ return err
+ }
+ var block strings.Builder
+ for _, e := range entries {
+ block.WriteString(e.Serialize())
+ }
+ if v.Passphrase == "" {
+ f, err := os.OpenFile(path, os.O_APPEND|os.O_WRONLY, 0o600) //nolint:gosec // Day path derives from the vault directory.
+ if err != nil {
+ return fmt.Errorf("open day file: %w", err)
+ }
+ defer func() { _ = f.Close() }()
+ if _, err := f.WriteString(block.String()); err != nil {
+ return fmt.Errorf("append entries: %w", err)
+ }
+ return nil
+ }
+ existing, err := v.readDayBytes(path)
+ if err != nil {
+ return err
+ }
+ return v.writeDayBytes(path, append(existing, []byte(block.String())...))
+}
+
+// ImportMarkers are the body lines that mark an entry as produced by an
+// importer rather than written by the person. They live here so every command
+// judging "did the person write this" shares one definition.
+var ImportMarkers = []string{"ICS-UID: ", "GIT-COMMIT: "}
+
+// Authored reports whether the entry was written by the person rather than
+// imported. The distinction is load-bearing: a backfilled vault holds thousands
+// of imported entries, and any feature that means to measure the person's own
+// writing must not count them.
+func (e Entry) Authored() bool {
+ for line := range strings.SplitSeq(e.Body, "\n") {
+ trimmed := strings.TrimSpace(line)
+ for _, m := range ImportMarkers {
+ if strings.HasPrefix(trimmed, m) {
+ return false
+ }
+ }
+ }
+ return true
+}
+
// Serialize renders the entry as the markdown block written to a day file.
// The block begins with a level-two header carrying the time and tags
// and ends with a trailing blank line so successive entries stay separated.
diff --git a/internal/vault/entry_test.go b/internal/vault/entry_test.go
index 680ca0e..5e860a2 100644
--- a/internal/vault/entry_test.go
+++ b/internal/vault/entry_test.go
@@ -82,3 +82,64 @@ func TestEntryHasTag(t *testing.T) {
})
}
}
+
+func TestAuthored(t *testing.T) {
+ t.Parallel()
+ tests := []struct {
+ Name string
+ Body string
+ Want bool
+ }{{ // Test 0: A plain written entry is authored.
+ Name: "handwritten", Body: "Thought about the roadmap today.", Want: true,
+ }, { // Test 1: A calendar import is not.
+ Name: "calendar", Body: "Standup (09:00 to 09:15)\nICS-UID: abc@example.com", Want: false,
+ }, { // Test 2: A commit import is not.
+ Name: "git", Body: "Fix the bug\nRepo: midden\nGIT-COMMIT: deadbeef", Want: false,
+ }, { // Test 3: Prose mentioning a marker mid-line stays authored.
+ Name: "prose mention", Body: "Wrote about how the ICS-UID: format works.", Want: true,
+ }, { // Test 4: An answer to an ask question is authored.
+ Name: "answer", Body: "He switched to baseball.\n\nIn answer to: something\nMIDDEN-ASKED: ended-abc", Want: true,
+ }}
+ for testNum, test := range tests {
+ t.Run(fmt.Sprintf("test %d %s", testNum, test.Name), func(t *testing.T) {
+ t.Parallel()
+ if got := (Entry{Body: test.Body}).Authored(); got != test.Want {
+ t.Errorf("want %v, got %v", test.Want, got)
+ }
+ })
+ }
+}
+
+func TestStreakCountsOnlyKeptEntries(t *testing.T) {
+ t.Parallel()
+ v, err := Open(t.TempDir())
+ if err != nil {
+ t.Fatalf("Open: %v", err)
+ }
+ today := time.Date(2026, time.August, 26, 12, 0, 0, 0, time.Local)
+ entries := []Entry{
+ // Imported events on today and yesterday, a real entry only yesterday.
+ {Time: today, Body: "Standup\nICS-UID: a@b"},
+ {Time: today.AddDate(0, 0, -1), Body: "Standup\nICS-UID: c@d"},
+ {Time: today.AddDate(0, 0, -1), Body: "I wrote this myself."},
+ }
+ if err := v.AppendAll(entries); err != nil {
+ t.Fatalf("AppendAll: %v", err)
+ }
+ all, err := v.Streak(today, nil)
+ if err != nil {
+ t.Fatalf("Streak: %v", err)
+ }
+ if all != 2 {
+ t.Errorf("want an unfiltered streak of 2, got %d", all)
+ }
+ authored, err := v.Streak(today, Entry.Authored)
+ if err != nil {
+ t.Fatalf("Streak: %v", err)
+ }
+ // Today holds only an imported event, so the authored streak is broken at
+ // today: appointments attended are not writing.
+ if authored != 0 {
+ t.Errorf("want an authored streak of 0, got %d", authored)
+ }
+}
diff --git a/internal/vault/recent_test.go b/internal/vault/recent_test.go
new file mode 100644
index 0000000..623e587
--- /dev/null
+++ b/internal/vault/recent_test.go
@@ -0,0 +1,139 @@
+package vault
+
+import (
+ "testing"
+ "time"
+
+ "github.com/google/go-cmp/cmp"
+)
+
+// scheduledFixture returns a vault holding past entries plus a calendar
+// appointment dated well into the future, which is what an imported calendar
+// puts in a vault.
+func scheduledFixture(t *testing.T) (*Vault, time.Time) {
+ t.Helper()
+ v, err := Open(t.TempDir())
+ if err != nil {
+ t.Fatalf("Open: %v", err)
+ }
+ now := time.Date(2026, time.August, 26, 12, 0, 0, 0, time.Local)
+ entries := []Entry{
+ {Time: now.AddDate(0, 0, -30), Body: "older"},
+ {Time: now.AddDate(0, 0, -1), Body: "yesterday"},
+ {Time: now.Add(-2 * time.Hour), Body: "this morning"},
+ {Time: now.AddDate(0, 7, 0), Body: "dentist next spring"},
+ }
+ if err := v.AppendAll(entries); err != nil {
+ t.Fatalf("AppendAll: %v", err)
+ }
+ return v, now
+}
+
+func TestRecentBeforeExcludesScheduledEntries(t *testing.T) {
+ t.Parallel()
+ v, now := scheduledFixture(t)
+ got, err := v.RecentBefore(2, now)
+ if err != nil {
+ t.Fatalf("RecentBefore: %v", err)
+ }
+ bodies := make([]string, len(got))
+ for i, e := range got {
+ bodies[i] = e.Body
+ }
+ // Without a cutoff the newest entry is an appointment that has not happened,
+ // so a vault backfilled in August reports next spring as the latest thing.
+ if diff := cmp.Diff([]string{"this morning", "yesterday"}, bodies); diff != "" {
+ t.Errorf("mismatch (-want +got):\n%s", diff)
+ }
+}
+
+func TestRecentBeforeZeroCutoffIncludesEverything(t *testing.T) {
+ t.Parallel()
+ v, _ := scheduledFixture(t)
+ got, err := v.RecentBefore(1, time.Time{})
+ if err != nil {
+ t.Fatalf("RecentBefore: %v", err)
+ }
+ if len(got) != 1 || got[0].Body != "dentist next spring" {
+ t.Errorf("want an open cutoff to reach the scheduled entry, got %v", got)
+ }
+}
+
+func TestRecentKeepsOpenBehavior(t *testing.T) {
+ t.Parallel()
+ v, _ := scheduledFixture(t)
+ got, err := v.Recent(1)
+ if err != nil {
+ t.Fatalf("Recent: %v", err)
+ }
+ if len(got) != 1 || got[0].Body != "dentist next spring" {
+ t.Errorf("want Recent to stay unbounded, got %v", got)
+ }
+}
+
+func TestComputeStatsSeparatesScheduledFromPast(t *testing.T) {
+ t.Parallel()
+ v, now := scheduledFixture(t)
+ s, err := v.ComputeStats(0, now)
+ if err != nil {
+ t.Fatalf("ComputeStats: %v", err)
+ }
+ if diff := cmp.Diff(1, s.Scheduled); diff != "" {
+ t.Errorf("scheduled count mismatch (-want +got):\n%s", diff)
+ }
+ // The span a person means is what has happened, not what is booked.
+ if !s.LastPast.Before(now) && !s.LastPast.Equal(now) {
+ t.Errorf("want the last past entry at or before now, got %s", s.LastPast)
+ }
+ if !s.LastEntry.After(now) {
+ t.Errorf("want the overall last entry to still reach the scheduled one, got %s", s.LastEntry)
+ }
+ if diff := cmp.Diff(4, s.Entries); diff != "" {
+ t.Errorf("want every entry counted (-want +got):\n%s", diff)
+ }
+}
+
+func TestComputeStatsZeroNowCountsNothingScheduled(t *testing.T) {
+ t.Parallel()
+ v, _ := scheduledFixture(t)
+ s, err := v.ComputeStats(0, time.Time{})
+ if err != nil {
+ t.Fatalf("ComputeStats: %v", err)
+ }
+ if s.Scheduled != 0 {
+ t.Errorf("want nothing scheduled without an observation time, got %d", s.Scheduled)
+ }
+}
+
+func TestComputeStatsCountsAuthoredSeparately(t *testing.T) {
+ t.Parallel()
+ v, err := Open(t.TempDir())
+ if err != nil {
+ t.Fatalf("Open: %v", err)
+ }
+ now := time.Date(2026, time.August, 26, 12, 0, 0, 0, time.Local)
+ entries := []Entry{
+ {Time: now.AddDate(0, 0, -10), Body: "Standup\nICS-UID: a@b"},
+ {Time: now.AddDate(0, 0, -9), Body: "Fix bug\nGIT-COMMIT: deadbeef"},
+ {Time: now.AddDate(0, 0, -8), Body: "I wrote this."},
+ {Time: now.AddDate(0, 0, -2), Body: "And this."},
+ // A scheduled import must count as neither past nor authored-latest.
+ {Time: now.AddDate(0, 1, 0), Body: "Dentist\nICS-UID: c@d"},
+ }
+ if err := v.AppendAll(entries); err != nil {
+ t.Fatalf("AppendAll: %v", err)
+ }
+ s, err := v.ComputeStats(0, now)
+ if err != nil {
+ t.Fatalf("ComputeStats: %v", err)
+ }
+ // This is the number the experiment turns on, so it must never drift into
+ // counting imports.
+ if diff := cmp.Diff(2, s.Authored); diff != "" {
+ t.Errorf("authored mismatch (-want +got):\n%s", diff)
+ }
+ want := now.AddDate(0, 0, -2)
+ if !s.LastAuthored.Equal(want) {
+ t.Errorf("want last authored %s, got %s", want, s.LastAuthored)
+ }
+}
diff --git a/internal/vault/search.go b/internal/vault/search.go
index dee794e..9cfd094 100644
--- a/internal/vault/search.go
+++ b/internal/vault/search.go
@@ -38,6 +38,17 @@ func (v *Vault) ReadRange(from, to time.Time) ([]Entry, error) {
// timestamp before selection because imports can append out of order; an
// entry's timestamp always falls on its file's date, so day order holds.
func (v *Vault) Recent(n int) ([]Entry, error) {
+ return v.RecentBefore(n, time.Time{})
+}
+
+// RecentBefore returns the most recent n entries at or before the cutoff,
+// newest first. A zero cutoff includes everything.
+//
+// The cutoff exists because a vault holding an imported calendar contains
+// appointments that have not happened yet. Without it "the most recent entry"
+// means the furthest one in the future, so a vault backfilled in August reports
+// next March's dentist appointment as the latest thing in the record.
+func (v *Vault) RecentBefore(n int, cutoff time.Time) ([]Entry, error) {
if n <= 0 {
return nil, fmt.Errorf("recent count must be positive")
}
@@ -47,12 +58,18 @@ func (v *Vault) Recent(n int) ([]Entry, error) {
}
var out []Entry
for i := len(days) - 1; i >= 0 && len(out) < n; i-- {
+ if !cutoff.IsZero() && days[i].After(cutoff) {
+ continue
+ }
entries, err := v.ReadDay(days[i])
if err != nil {
return nil, err
}
sort.SliceStable(entries, func(a, b int) bool { return entries[a].Time.Before(entries[b].Time) })
for j := len(entries) - 1; j >= 0 && len(out) < n; j-- {
+ if !cutoff.IsZero() && entries[j].Time.After(cutoff) {
+ continue
+ }
out = append(out, entries[j])
}
}
diff --git a/internal/vault/stats.go b/internal/vault/stats.go
index 8114ee1..0a772ea 100644
--- a/internal/vault/stats.go
+++ b/internal/vault/stats.go
@@ -2,9 +2,10 @@ package vault
import (
"fmt"
- "sort"
"strings"
"time"
+
+ "github.com/dcadolph/midden/internal/util"
)
// Stats is a summary of the vault contents.
@@ -22,20 +23,30 @@ type Stats struct {
// LastEntry is the timestamp of the latest entry, or the zero time when the vault is empty.
LastEntry time.Time `json:"last_entry"`
// TopTags is the tag histogram ordered by descending count then label.
- TopTags []TagCount `json:"top_tags,omitempty"`
-}
-
-// TagCount pairs a tag label with the number of entries that carry it.
-type TagCount struct {
- // Tag is the tag label without the leading hash character.
- Tag string `json:"tag"`
- // Count is the number of entries carrying the tag.
- Count int `json:"count"`
+ TopTags []util.TagCount `json:"top_tags,omitempty"`
+ // Scheduled is the number of entries dated after the observation time. An
+ // imported calendar carries appointments that have not happened yet, and
+ // counting them among what the record holds overstates it.
+ Scheduled int `json:"scheduled"`
+ // LastPast is the latest entry at or before the observation time, which is
+ // what a person means by the most recent thing in the record.
+ LastPast time.Time `json:"last_past"`
+ // Authored is the number of entries the person wrote rather than imported.
+ // This is the number the whole experiment turns on: imported history proves
+ // the tool can hold a life, and only this count proves the person has
+ // started giving it the part no import can reach.
+ Authored int `json:"authored"`
+ // LastAuthored is the most recent authored entry, or the zero time when
+ // nothing has been written yet.
+ LastAuthored time.Time `json:"last_authored,omitempty"`
}
// ComputeStats walks every entry once and assembles the summary.
-// TopTags is capped at the given limit; pass a non-positive limit to include every tag.
-func (v *Vault) ComputeStats(topTagLimit int) (Stats, error) {
+// TopTags is capped at the given limit; pass a non-positive limit to include
+// every tag. Entries dated after now are counted separately as scheduled rather
+// than folded into the record's span, so an imported calendar does not make the
+// vault appear to run into next year.
+func (v *Vault) ComputeStats(topTagLimit int, now time.Time) (Stats, error) {
var s Stats
days, err := v.ListDays()
if err != nil {
@@ -60,19 +71,31 @@ func (v *Vault) ComputeStats(topTagLimit int) (Stats, error) {
if e.Time.After(s.LastEntry) {
s.LastEntry = e.Time
}
+ switch {
+ case !now.IsZero() && e.Time.After(now):
+ s.Scheduled++
+ case e.Time.After(s.LastPast):
+ s.LastPast = e.Time
+ }
+ if e.Authored() {
+ s.Authored++
+ if e.Time.After(s.LastAuthored) && (now.IsZero() || !e.Time.After(now)) {
+ s.LastAuthored = e.Time
+ }
+ }
for _, t := range e.Tags {
tagCounts[strings.ToLower(t)]++
}
}
}
s.Tags = len(tagCounts)
- s.TopTags = sortedTagCounts(tagCounts, topTagLimit)
+ s.TopTags = util.SortedCounts(tagCounts, topTagLimit)
return s, nil
}
// TagCounts returns every distinct tag with its entry count, ordered by descending count then label.
// Pass a non-positive limit to include every tag.
-func (v *Vault) TagCounts(limit int) ([]TagCount, error) {
+func (v *Vault) TagCounts(limit int) ([]util.TagCount, error) {
days, err := v.ListDays()
if err != nil {
return nil, err
@@ -89,12 +112,17 @@ func (v *Vault) TagCounts(limit int) ([]TagCount, error) {
}
}
}
- return sortedTagCounts(tagCounts, limit), nil
+ return util.SortedCounts(tagCounts, limit), nil
}
-// Streak returns the number of consecutive days ending today on which at least one entry was written.
-// A day with no entry breaks the streak, including today.
-func (v *Vault) Streak(today time.Time) (int, error) {
+// Streak returns the number of consecutive days ending today on which at least
+// one entry satisfying keep was written. A day with no such entry breaks the
+// streak, including today. A nil keep counts every entry.
+//
+// The filter exists because a backfilled vault has entries on thousands of days
+// the person never wrote a word: a streak counted over imported calendar events
+// congratulates them for appointments they merely attended.
+func (v *Vault) Streak(today time.Time, keep func(Entry) bool) (int, error) {
today = dayStart(today)
days, err := v.ListDays()
if err != nil {
@@ -106,8 +134,11 @@ func (v *Vault) Streak(today time.Time) (int, error) {
if err != nil {
return 0, err
}
- if len(entries) > 0 {
- have[dayKey(d)] = true
+ for _, e := range entries {
+ if keep == nil || keep(e) {
+ have[dayKey(d)] = true
+ break
+ }
}
}
streak := 0
@@ -152,27 +183,3 @@ func countWords(s string) int {
func dayKey(t time.Time) string {
return t.Format("2006-01-02")
}
-
-// sortedTagCounts returns the histogram as a slice ordered by descending count and ascending label.
-// A non-positive limit returns every entry.
-func sortedTagCounts(counts map[string]int, limit int) []TagCount {
- out := make([]TagCount, 0, len(counts))
- for t, n := range counts {
- out = append(out, TagCount{Tag: t, Count: n})
- }
- sortByCountDescThenLabel(out)
- if limit > 0 && len(out) > limit {
- out = out[:limit]
- }
- return out
-}
-
-// sortByCountDescThenLabel orders the slice in place by descending count then ascending label.
-func sortByCountDescThenLabel(out []TagCount) {
- sort.Slice(out, func(i, j int) bool {
- if out[i].Count != out[j].Count {
- return out[i].Count > out[j].Count
- }
- return out[i].Tag < out[j].Tag
- })
-}
diff --git a/internal/vault/stats_test.go b/internal/vault/stats_test.go
index 343c56c..ae7512d 100644
--- a/internal/vault/stats_test.go
+++ b/internal/vault/stats_test.go
@@ -5,12 +5,14 @@ import (
"time"
"github.com/google/go-cmp/cmp"
+
+ "github.com/dcadolph/midden/internal/util"
)
func TestComputeStats(t *testing.T) {
t.Parallel()
v := seedFixture(t)
- got, err := v.ComputeStats(0)
+ got, err := v.ComputeStats(0, time.Time{})
if err != nil {
t.Fatalf("ComputeStats: %v", err)
}
@@ -21,7 +23,13 @@ func TestComputeStats(t *testing.T) {
Tags: 2,
FirstEntry: time.Date(2026, 6, 16, 9, 14, 23, 0, time.Local),
LastEntry: time.Date(2026, 6, 18, 10, 0, 0, 0, time.Local),
- TopTags: []TagCount{
+ // With no observation time nothing counts as scheduled, so the last past
+ // entry is simply the last entry.
+ LastPast: time.Date(2026, 6, 18, 10, 0, 0, 0, time.Local),
+ // The fixture's entries are all hand-written, so every one is authored.
+ Authored: 3,
+ LastAuthored: time.Date(2026, 6, 18, 10, 0, 0, 0, time.Local),
+ TopTags: []util.TagCount{
{Tag: "life", Count: 1},
{Tag: "project", Count: 1},
},
@@ -61,7 +69,7 @@ func TestStreakCountsConsecutiveDays(t *testing.T) {
t.Fatalf("Append: %v", err)
}
}
- streak, err := v.Streak(today)
+ streak, err := v.Streak(today, nil)
if err != nil {
t.Fatalf("Streak: %v", err)
}
@@ -80,7 +88,7 @@ func TestStreakIsZeroWhenTodayMissing(t *testing.T) {
if err := v.Append(Entry{Time: today.AddDate(0, 0, -1), Body: "yesterday"}); err != nil {
t.Fatalf("Append: %v", err)
}
- streak, err := v.Streak(today)
+ streak, err := v.Streak(today, nil)
if err != nil {
t.Fatalf("Streak: %v", err)
}
diff --git a/internal/weave/gap.go b/internal/weave/gap.go
new file mode 100644
index 0000000..7fe1172
--- /dev/null
+++ b/internal/weave/gap.go
@@ -0,0 +1,228 @@
+package weave
+
+import (
+ "sort"
+ "time"
+
+ "github.com/dcadolph/midden/internal/vault"
+)
+
+// Gap is a stretch where the record went quiet relative to the activity around
+// it. A thread ending is one commitment stopping; a gap is the record itself
+// falling silent, which is a different and usually larger thing. It is also
+// invisible from the inside: a person notices that a class ended, but never
+// that a whole span of years went unrecorded.
+type Gap struct {
+ // Source is the source tag whose record fell silent, or empty when the
+ // silence spans the whole record. The distinction matters because one loud
+ // source can flood the months where another went quiet: commits pouring in
+ // during years the calendar recorded nothing would otherwise hide exactly
+ // the silence worth asking about.
+ Source string
+ // From is the first day of the first quiet month.
+ From time.Time
+ // To is the last day of the last quiet month.
+ To time.Time
+ // Months is how long the silence ran.
+ Months int
+ // Entries is how many entries fall inside it.
+ Entries int
+ // Before and After are entries per month in the active stretches on either
+ // side, which is what makes the quiet legible as a departure.
+ Before float64
+ After float64
+}
+
+// GapOptions tune silence detection.
+type GapOptions struct {
+ // Now is the observation date. Months after it are ignored, since a record
+ // cannot be silent about a future that has not happened.
+ Now time.Time
+ // MinMonths is the shortest silence worth reporting.
+ MinMonths int
+ // QuietFraction is the share of the record's usual monthly volume at or
+ // below which a month counts as quiet.
+ QuietFraction float64
+}
+
+// DefaultGapOptions returns silence settings suited to a personal record.
+func DefaultGapOptions(now time.Time) GapOptions {
+ return GapOptions{Now: now, MinMonths: 6, QuietFraction: 0.15}
+}
+
+// GapsBySource finds interior silences in the whole record and inside each
+// source separately, longest first. A silence found in the whole record is not
+// repeated per source.
+func GapsBySource(entries []vault.Entry, sources []string, opts GapOptions) []Gap {
+ out := Gaps(entries, opts)
+ covered := func(g Gap) bool {
+ for _, w := range out {
+ if w.Source == "" && !g.From.Before(w.From) && !g.To.After(w.To) {
+ return true
+ }
+ }
+ return false
+ }
+ for _, src := range sources {
+ var subset []vault.Entry
+ for _, e := range entries {
+ if sourceOf(e, []string{src}) != "" {
+ subset = append(subset, e)
+ }
+ }
+ for _, g := range Gaps(subset, opts) {
+ g.Source = src
+ if !covered(g) {
+ out = append(out, g)
+ }
+ }
+ }
+ sort.Slice(out, func(a, b int) bool {
+ if out[a].Months != out[b].Months {
+ return out[a].Months > out[b].Months
+ }
+ return out[a].From.Before(out[b].From)
+ })
+ return out
+}
+
+// Gaps finds the interior silences in a record, longest first.
+//
+// Only interior silences count. A record is quiet before it starts and after it
+// ends by definition, and reporting those as holes would say nothing about the
+// life, so a gap is required to have activity on both sides.
+func Gaps(entries []vault.Entry, opts GapOptions) []Gap {
+ if opts.MinMonths < 1 {
+ opts.MinMonths = 1
+ }
+ months, first, last := monthCounts(entries, opts.Now)
+ if len(months) == 0 || !first.Before(last) {
+ return nil
+ }
+ series := monthSeries(first, last)
+ if len(series) < opts.MinMonths+2 {
+ return nil
+ }
+ threshold := quietThreshold(months, series, opts.QuietFraction)
+
+ var out []Gap
+ i := 0
+ for i < len(series) {
+ if months[key(series[i])] > threshold {
+ i++
+ continue
+ }
+ j := i
+ for j < len(series) && months[key(series[j])] <= threshold {
+ j++
+ }
+ // Interior only: something has to come before and after the silence.
+ if i > 0 && j < len(series) && j-i >= opts.MinMonths {
+ out = append(out, buildGap(series, months, i, j))
+ }
+ i = j
+ }
+ sort.Slice(out, func(a, b int) bool {
+ if out[a].Months != out[b].Months {
+ return out[a].Months > out[b].Months
+ }
+ return out[a].From.Before(out[b].From)
+ })
+ return out
+}
+
+// buildGap assembles the gap spanning series[i:j].
+func buildGap(series []time.Time, months map[string]int, i, j int) Gap {
+ start := series[i]
+ end := series[j-1].AddDate(0, 1, -1)
+ inside := 0
+ for k := i; k < j; k++ {
+ inside += months[key(series[k])]
+ }
+ return Gap{
+ From: start,
+ To: end,
+ Months: j - i,
+ Entries: inside,
+ Before: rate(series, months, 0, i),
+ After: rate(series, months, j, len(series)),
+ }
+}
+
+// rate returns entries per month across series[from:to].
+func rate(series []time.Time, months map[string]int, from, to int) float64 {
+ if to <= from {
+ return 0
+ }
+ total := 0
+ for k := from; k < to; k++ {
+ total += months[key(series[k])]
+ }
+ return float64(total) / float64(to-from)
+}
+
+// quietThreshold returns the monthly count at or below which a month reads as
+// quiet. It is a fraction of the median active month rather than of the mean,
+// because a single explosive month would otherwise drag the bar high enough to
+// call ordinary months silent.
+func quietThreshold(months map[string]int, series []time.Time, fraction float64) int {
+ var active []int
+ for _, m := range series {
+ if n := months[key(m)]; n > 0 {
+ active = append(active, n)
+ }
+ }
+ if len(active) == 0 {
+ return 0
+ }
+ sort.Ints(active)
+ median := active[len(active)/2]
+ if fraction <= 0 {
+ return 0
+ }
+ t := int(float64(median) * fraction)
+ if t < 1 {
+ t = 1
+ }
+ return t
+}
+
+// monthCounts buckets entries by month and returns the bounding months.
+func monthCounts(entries []vault.Entry, now time.Time) (map[string]int, time.Time, time.Time) {
+ out := map[string]int{}
+ var first, last time.Time
+ for _, e := range entries {
+ if !now.IsZero() && e.Time.After(now) {
+ continue
+ }
+ m := monthOf(e.Time)
+ out[key(m)]++
+ if first.IsZero() || m.Before(first) {
+ first = m
+ }
+ if m.After(last) {
+ last = m
+ }
+ }
+ return out, first, last
+}
+
+// monthSeries returns every month from first to last inclusive, including the
+// empty ones, which are the whole point.
+func monthSeries(first, last time.Time) []time.Time {
+ var out []time.Time
+ for m := first; !m.After(last); m = m.AddDate(0, 1, 0) {
+ out = append(out, m)
+ }
+ return out
+}
+
+// monthOf truncates a timestamp to the first day of its month.
+func monthOf(t time.Time) time.Time {
+ return time.Date(t.Year(), t.Month(), 1, 0, 0, 0, 0, t.Location())
+}
+
+// key renders a month as its bucket key.
+func key(m time.Time) string {
+ return m.Format(monthLayout)
+}
diff --git a/internal/weave/gap_test.go b/internal/weave/gap_test.go
new file mode 100644
index 0000000..9ac92d3
--- /dev/null
+++ b/internal/weave/gap_test.go
@@ -0,0 +1,194 @@
+package weave
+
+import (
+ "testing"
+ "time"
+
+ "github.com/dcadolph/midden/internal/vault"
+)
+
+// monthly builds n entries spread through the given month.
+func monthly(year int, month time.Month, n int) []vault.Entry {
+ out := make([]vault.Entry, 0, n)
+ for i := range n {
+ day := 1 + i%27
+ out = append(out, vault.Entry{
+ Time: time.Date(year, month, day, 12, 0, 0, 0, time.Local),
+ Body: "entry",
+ })
+ }
+ return out
+}
+
+// span builds entries across every month from one date to another.
+func span(fromYear int, fromMonth time.Month, months, perMonth int) []vault.Entry {
+ var out []vault.Entry
+ cur := time.Date(fromYear, fromMonth, 1, 0, 0, 0, 0, time.Local)
+ for range months {
+ out = append(out, monthly(cur.Year(), cur.Month(), perMonth)...)
+ cur = cur.AddDate(0, 1, 0)
+ }
+ return out
+}
+
+func TestGapsFindAnInteriorSilence(t *testing.T) {
+ t.Parallel()
+ var entries []vault.Entry
+ entries = append(entries, span(2015, time.January, 12, 20)...) // active
+ // 2016 through 2017 silent
+ entries = append(entries, span(2018, time.January, 12, 20)...) // active again
+ got := Gaps(entries, DefaultGapOptions(time.Date(2019, time.January, 1, 0, 0, 0, 0, time.Local)))
+ if len(got) != 1 {
+ t.Fatalf("want one silence, got %d: %+v", len(got), got)
+ }
+ if got[0].Months != 24 {
+ t.Errorf("want a 24 month silence, got %d", got[0].Months)
+ }
+ if got[0].Entries != 0 {
+ t.Errorf("want nothing inside the silence, got %d", got[0].Entries)
+ }
+ if got[0].Before <= 0 || got[0].After <= 0 {
+ t.Errorf("want activity measured on both sides, got %.1f and %.1f", got[0].Before, got[0].After)
+ }
+}
+
+func TestGapsIgnoreTheEdgesOfARecord(t *testing.T) {
+ t.Parallel()
+ // A record is quiet before it starts and after it ends by definition.
+ // Reporting those as holes would say nothing about the life.
+ entries := span(2020, time.January, 12, 20)
+ got := Gaps(entries, DefaultGapOptions(time.Date(2026, time.January, 1, 0, 0, 0, 0, time.Local)))
+ if len(got) != 0 {
+ t.Errorf("want no silence for a record that simply ended, got %+v", got)
+ }
+}
+
+func TestGapsHonorMinMonths(t *testing.T) {
+ t.Parallel()
+ var entries []vault.Entry
+ entries = append(entries, span(2020, time.January, 6, 20)...)
+ // A three month lull, shorter than the floor.
+ entries = append(entries, span(2020, time.October, 6, 20)...)
+ opts := DefaultGapOptions(time.Date(2021, time.June, 1, 0, 0, 0, 0, time.Local))
+ if got := Gaps(entries, opts); len(got) != 0 {
+ t.Errorf("want a short lull ignored, got %+v", got)
+ }
+ opts.MinMonths = 3
+ if got := Gaps(entries, opts); len(got) != 1 {
+ t.Errorf("want the lull found at a lower floor, got %d", len(got))
+ }
+}
+
+func TestGapsToleratesSparseMonthsInsideASilence(t *testing.T) {
+ t.Parallel()
+ var entries []vault.Entry
+ entries = append(entries, span(2015, time.January, 12, 40)...)
+ // Two stray entries during the quiet stretch, which is what a real silence
+ // looks like: not empty, just nearly so.
+ entries = append(entries, monthly(2016, time.June, 1)...)
+ entries = append(entries, monthly(2017, time.March, 1)...)
+ entries = append(entries, span(2018, time.January, 12, 40)...)
+ got := Gaps(entries, DefaultGapOptions(time.Date(2019, time.January, 1, 0, 0, 0, 0, time.Local)))
+ if len(got) != 1 {
+ t.Fatalf("want the near-silence still found, got %d: %+v", len(got), got)
+ }
+ if got[0].Entries != 2 {
+ t.Errorf("want the stray entries counted inside it, got %d", got[0].Entries)
+ }
+}
+
+func TestGapsIgnoreFutureMonths(t *testing.T) {
+ t.Parallel()
+ var entries []vault.Entry
+ entries = append(entries, span(2025, time.January, 6, 20)...)
+ // A scheduled appointment far ahead must not open a silence between now and it.
+ entries = append(entries, monthly(2028, time.March, 1)...)
+ got := Gaps(entries, DefaultGapOptions(time.Date(2026, time.August, 26, 0, 0, 0, 0, time.Local)))
+ if len(got) != 0 {
+ t.Errorf("want future months excluded, got %+v", got)
+ }
+}
+
+func TestGapsEmptyRecord(t *testing.T) {
+ t.Parallel()
+ if got := Gaps(nil, DefaultGapOptions(time.Now())); len(got) != 0 {
+ t.Errorf("want nothing from an empty record, got %+v", got)
+ }
+}
+
+func TestGapsOrderLongestFirst(t *testing.T) {
+ t.Parallel()
+ var entries []vault.Entry
+ entries = append(entries, span(2010, time.January, 6, 20)...)
+ entries = append(entries, span(2011, time.July, 6, 20)...) // after a 12 month gap
+ entries = append(entries, span(2014, time.July, 6, 20)...) // after a 30 month gap
+ got := Gaps(entries, DefaultGapOptions(time.Date(2015, time.June, 1, 0, 0, 0, 0, time.Local)))
+ if len(got) != 2 {
+ t.Fatalf("want two silences, got %d", len(got))
+ }
+ if got[0].Months < got[1].Months {
+ t.Errorf("want the longer silence first, got %d then %d", got[0].Months, got[1].Months)
+ }
+}
+
+func TestGapsBySourceFindsASilenceOneSourceMasks(t *testing.T) {
+ t.Parallel()
+ var entries []vault.Entry
+ // The calendar goes quiet for two years while commits flood the same months.
+ for _, e := range span(2022, time.January, 12, 8) {
+ e.Tags = []string{"calendar"}
+ entries = append(entries, e)
+ }
+ for _, e := range span(2025, time.January, 12, 8) {
+ e.Tags = []string{"calendar"}
+ entries = append(entries, e)
+ }
+ for _, e := range span(2022, time.January, 48, 30) {
+ e.Tags = []string{"git"}
+ entries = append(entries, e)
+ }
+ now := time.Date(2026, time.January, 1, 0, 0, 0, 0, time.Local)
+ // The whole record never goes quiet, which is exactly how one loud source
+ // hides the silence in another.
+ if got := Gaps(entries, DefaultGapOptions(now)); len(got) != 0 {
+ t.Fatalf("precondition: want no whole-record silence, got %+v", got)
+ }
+ got := GapsBySource(entries, []string{"calendar", "git"}, DefaultGapOptions(now))
+ if len(got) != 1 {
+ t.Fatalf("want the masked calendar silence found, got %d: %+v", len(got), got)
+ }
+ if got[0].Source != "calendar" {
+ t.Errorf("want the silence attributed to the calendar, got %q", got[0].Source)
+ }
+ if got[0].Months < 20 {
+ t.Errorf("want roughly two years of silence, got %d months", got[0].Months)
+ }
+}
+
+func TestGapsBySourceDoesNotRepeatAWholeRecordSilence(t *testing.T) {
+ t.Parallel()
+ var entries []vault.Entry
+ // Every source is quiet over the same stretch, so the whole-record gap
+ // already covers it and per-source copies would ask the same question twice.
+ for _, e := range span(2020, time.January, 12, 10) {
+ e.Tags = []string{"calendar"}
+ entries = append(entries, e)
+ }
+ for _, e := range span(2023, time.January, 12, 10) {
+ e.Tags = []string{"calendar"}
+ entries = append(entries, e)
+ }
+ now := time.Date(2024, time.January, 1, 0, 0, 0, 0, time.Local)
+ got := GapsBySource(entries, []string{"calendar"}, DefaultGapOptions(now))
+ whole, perSource := 0, 0
+ for _, g := range got {
+ if g.Source == "" {
+ whole++
+ } else {
+ perSource++
+ }
+ }
+ if whole != 1 || perSource != 0 {
+ t.Errorf("want one whole-record silence and no per-source copy, got %d and %d", whole, perSource)
+ }
+}
diff --git a/internal/weave/handoff.go b/internal/weave/handoff.go
new file mode 100644
index 0000000..06a3178
--- /dev/null
+++ b/internal/weave/handoff.go
@@ -0,0 +1,118 @@
+package weave
+
+import (
+ "sort"
+ "time"
+)
+
+// Handoff pairs a thread that ended with one that began soon after. A person
+// remembers the activities themselves but rarely the succession between them,
+// and the succession is where the story is: a thing stopped, and something took
+// its place a few weeks later.
+type Handoff struct {
+ // From is the thread that ended.
+ From Thread
+ // To is the thread that began after it.
+ To Thread
+ // GapDays is how long passed between the last of one and the first of the other.
+ GapDays int
+}
+
+// HandoffOptions tune succession detection.
+type HandoffOptions struct {
+ // Window is the most days that may pass between an ending and a beginning
+ // for the two to count as a succession.
+ Window int
+ // SameSubject reports whether two threads concern the same subject. A
+ // handoff is only meaningful within one life, so a child dropping an
+ // activity should not pair with a sibling starting one. Nil pairs everything.
+ SameSubject func(a, b Thread) bool
+}
+
+// DefaultHandoffOptions returns succession settings suited to a personal record.
+// Successions are held to a shared subject, because one person dropping an
+// activity has nothing to do with another person starting one, and pairing them
+// would manufacture a story the record does not contain.
+func DefaultHandoffOptions() HandoffOptions {
+ return HandoffOptions{Window: 120, SameSubject: SharesSubject}
+}
+
+// SharesSubject reports whether two threads name a common subject, judged by a
+// word they both carry that is distinctive enough to be a person or activity.
+func SharesSubject(a, b Thread) bool {
+ aw, bw := wordSet(a.Key), wordSet(b.Key)
+ for w := range aw {
+ if bw[w] && len(w) > 2 {
+ return true
+ }
+ }
+ return false
+}
+
+// Handoffs finds successions among the given threads, ordered by the weight of
+// the thread that ended, so the largest losses come first.
+func Handoffs(threads []Thread, opts HandoffOptions) []Handoff {
+ if opts.Window <= 0 {
+ opts.Window = 120
+ }
+ var out []Handoff
+ for _, from := range threads {
+ if from.Status != Ended {
+ continue
+ }
+ for _, to := range threads {
+ if to.Key == from.Key || to.First.Before(from.Last) {
+ continue
+ }
+ gap := daysBetween(from.Last, to.First)
+ if gap > opts.Window {
+ continue
+ }
+ if opts.SameSubject != nil && !opts.SameSubject(from, to) {
+ continue
+ }
+ out = append(out, Handoff{From: from, To: to, GapDays: gap})
+ }
+ }
+ sort.Slice(out, func(i, j int) bool {
+ if out[i].From.Weight() != out[j].From.Weight() {
+ return out[i].From.Weight() > out[j].From.Weight()
+ }
+ return out[i].GapDays < out[j].GapDays
+ })
+ return out
+}
+
+// Milestone is a dated turning point in the record.
+type Milestone struct {
+ // When the turning point happened.
+ When time.Time
+ // Kind is what happened, either "ended" or "began".
+ Kind string
+ // Thread is the thread that turned.
+ Thread Thread
+}
+
+// Milestones returns the endings and beginnings among the given threads in
+// chronological order, which is the order a life is told in.
+func Milestones(threads []Thread) []Milestone {
+ var out []Milestone
+ for _, t := range threads {
+ switch t.Status {
+ case Ended:
+ out = append(out, Milestone{When: t.Last, Kind: "ended", Thread: t})
+ case Emerging:
+ out = append(out, Milestone{When: t.First, Kind: "began", Thread: t})
+ case Ongoing, Dormant:
+ // Neither is a turning point: one is still running, the other is
+ // between seasons and has not turned anywhere.
+ }
+ }
+ sort.Slice(out, func(i, j int) bool {
+ if !out[i].When.Equal(out[j].When) {
+ return out[i].When.Before(out[j].When)
+ }
+ return out[i].Thread.Weight() > out[j].Thread.Weight()
+ })
+ return out
+}
diff --git a/internal/weave/handoff_test.go b/internal/weave/handoff_test.go
new file mode 100644
index 0000000..679a4df
--- /dev/null
+++ b/internal/weave/handoff_test.go
@@ -0,0 +1,150 @@
+package weave
+
+import (
+ "testing"
+ "time"
+
+ "github.com/google/go-cmp/cmp"
+
+ "github.com/dcadolph/midden/internal/vault"
+)
+
+// tagged builds an entry carrying a source tag.
+func tagged(year int, month time.Month, day int, tag, body string) vault.Entry {
+ return vault.Entry{
+ Time: time.Date(year, month, day, 12, 0, 0, 0, time.Local),
+ Tags: []string{tag},
+ Body: body,
+ }
+}
+
+func TestHandoffsPairSuccessionsWithinOneSubject(t *testing.T) {
+ t.Parallel()
+ var entries []vault.Entry
+ // One child drops an activity in October and picks up another that December.
+ entries = append(entries, weekly(2024, time.January, 6, 40, "William- martial arts")...)
+ entries = append(entries, weekly(2024, time.December, 1, 10, "William- baseball practice")...)
+ // A sibling starts something in the same window, which must not be paired.
+ entries = append(entries, weekly(2024, time.December, 2, 10, "Hannah- choir")...)
+
+ threads := Threads(entries, DefaultOptions(observed))
+ got := Handoffs(threads, DefaultHandoffOptions())
+ if len(got) == 0 {
+ t.Fatal("want at least one succession")
+ }
+ for _, h := range got {
+ if h.From.Label == "William- martial arts" && h.To.Label == "Hannah- choir" {
+ t.Error("want siblings kept apart: one child stopping has nothing to do with another starting")
+ }
+ }
+ found := false
+ for _, h := range got {
+ if h.From.Label == "William- martial arts" && h.To.Label == "William- baseball practice" {
+ found = true
+ if h.GapDays <= 0 {
+ t.Errorf("want a positive gap between the ending and the beginning, got %d", h.GapDays)
+ }
+ }
+ }
+ if !found {
+ t.Error("want the succession within one child's activities")
+ }
+}
+
+func TestHandoffsRespectTheWindow(t *testing.T) {
+ t.Parallel()
+ var entries []vault.Entry
+ entries = append(entries, weekly(2023, time.January, 5, 20, "William- piano")...)
+ // Begins years later, far outside any plausible succession.
+ entries = append(entries, weekly(2026, time.June, 4, 10, "William- baseball")...)
+ threads := Threads(entries, DefaultOptions(observed))
+ if got := Handoffs(threads, HandoffOptions{Window: 60, SameSubject: SharesSubject}); len(got) != 0 {
+ t.Errorf("want nothing paired across a multi-year gap, got %d", len(got))
+ }
+}
+
+func TestSharesSubject(t *testing.T) {
+ t.Parallel()
+ a := Thread{Key: "arts martial william"}
+ b := Thread{Key: "baseball william"}
+ c := Thread{Key: "choir hannah"}
+ if !SharesSubject(a, b) {
+ t.Error("want threads naming the same child to share a subject")
+ }
+ if SharesSubject(a, c) {
+ t.Error("want threads naming different children to be unrelated")
+ }
+}
+
+func TestMilestonesAreChronological(t *testing.T) {
+ t.Parallel()
+ var entries []vault.Entry
+ entries = append(entries, weekly(2024, time.January, 6, 30, "William- martial arts")...)
+ entries = append(entries, weekly(2026, time.July, 1, 8, "William- baseball")...)
+ got := Milestones(Threads(entries, DefaultOptions(observed)))
+ if len(got) < 2 {
+ t.Fatalf("want an ending and a beginning, got %d", len(got))
+ }
+ for i := 1; i < len(got); i++ {
+ if got[i].When.Before(got[i-1].When) {
+ t.Errorf("milestones out of order at %d: %s before %s", i, got[i].When, got[i-1].When)
+ }
+ }
+}
+
+func TestRhythmCountsBySourcePerMonth(t *testing.T) {
+ t.Parallel()
+ entries := []vault.Entry{
+ tagged(2026, time.July, 1, "calendar", "Dentist"),
+ tagged(2026, time.July, 2, "calendar", "Piano"),
+ tagged(2026, time.July, 3, "git", "Fix a bug"),
+ tagged(2026, time.August, 1, "git", "Ship it"),
+ }
+ got := Rhythm(entries, []string{"calendar", "git"}, observed)
+ if len(got) != 2 {
+ t.Fatalf("want two months, got %d", len(got))
+ }
+ if diff := cmp.Diff("2026-07", got[0].Month); diff != "" {
+ t.Errorf("months out of order (-want +got):\n%s", diff)
+ }
+ if got[0].Counts["calendar"] != 2 || got[0].Counts["git"] != 1 {
+ t.Errorf("July counts wrong: %v", got[0].Counts)
+ }
+ if got[0].Total != 3 {
+ t.Errorf("want a July total of 3, got %d", got[0].Total)
+ }
+}
+
+func TestOverlapsFindDaysWhereSourcesMeet(t *testing.T) {
+ t.Parallel()
+ entries := []vault.Entry{
+ tagged(2026, time.July, 30, "calendar", "Dad- off work/ vacation"),
+ tagged(2026, time.July, 30, "git", "Trim projects"),
+ tagged(2026, time.July, 30, "git", "Alphabetize projects"),
+ // A day carrying only one source is not a crossing.
+ tagged(2026, time.July, 31, "git", "Ship it"),
+ }
+ got := Overlaps(entries, []string{"calendar", "git"}, observed)
+ if len(got) != 1 {
+ t.Fatalf("want exactly the one day both sources recorded, got %d", len(got))
+ }
+ if got[0].Counts["git"] != 2 || got[0].Counts["calendar"] != 1 {
+ t.Errorf("counts wrong: %v", got[0].Counts)
+ }
+ // The headline is what makes a crossing legible: neither source alone says
+ // the day was spent working through a vacation.
+ if got[0].Headlines["calendar"] != "Dad- off work/ vacation" {
+ t.Errorf("want the calendar headline preserved, got %q", got[0].Headlines["calendar"])
+ }
+}
+
+func TestOverlapsIgnoreFutureDays(t *testing.T) {
+ t.Parallel()
+ entries := []vault.Entry{
+ tagged(2027, time.March, 12, "calendar", "Future appointment"),
+ tagged(2027, time.March, 12, "git", "Future commit"),
+ }
+ if got := Overlaps(entries, []string{"calendar", "git"}, observed); len(got) != 0 {
+ t.Errorf("want future days excluded, got %d", len(got))
+ }
+}
diff --git a/internal/weave/rhythm.go b/internal/weave/rhythm.go
new file mode 100644
index 0000000..2b4bb9c
--- /dev/null
+++ b/internal/weave/rhythm.go
@@ -0,0 +1,143 @@
+package weave
+
+import (
+ "sort"
+ "strings"
+ "time"
+
+ "github.com/dcadolph/midden/internal/vault"
+)
+
+// monthLayout renders a calendar month as YYYY-MM.
+const monthLayout = "2006-01"
+
+// Period is one calendar month with a count per source.
+type Period struct {
+ // Month is the calendar month in YYYY-MM form.
+ Month string
+ // Counts holds entries per source tag inside the month.
+ Counts map[string]int
+ // Total is the number of entries in the month across every source.
+ Total int
+}
+
+// Rhythm returns the record month by month, counted per source. A person knows
+// what they did in a given month but not how the balance between the parts of
+// their life shifted across years, because that comparison spans more time than
+// memory holds at once.
+func Rhythm(entries []vault.Entry, sources []string, now time.Time) []Period {
+ byMonth := map[string]map[string]int{}
+ for _, e := range entries {
+ if !now.IsZero() && e.Time.After(now) {
+ continue
+ }
+ src := sourceOf(e, sources)
+ if src == "" {
+ continue
+ }
+ m := e.Time.Format(monthLayout)
+ if byMonth[m] == nil {
+ byMonth[m] = map[string]int{}
+ }
+ byMonth[m][src]++
+ }
+ out := make([]Period, 0, len(byMonth))
+ for m, counts := range byMonth {
+ total := 0
+ for _, n := range counts {
+ total += n
+ }
+ out = append(out, Period{Month: m, Counts: counts, Total: total})
+ }
+ // The layout is zero-padded, so lexical order is calendar order.
+ sort.Slice(out, func(i, j int) bool { return out[i].Month < out[j].Month })
+ return out
+}
+
+// sourceOf returns which of the given source tags an entry carries, or empty
+// when it carries none. Sources are checked in order, so the caller controls
+// precedence when an entry carries more than one.
+func sourceOf(e vault.Entry, sources []string) string {
+ for _, want := range sources {
+ for _, t := range e.Tags {
+ if strings.EqualFold(t, want) {
+ return want
+ }
+ }
+ }
+ return ""
+}
+
+// Overlap is a day on which more than one source recorded something. A single
+// source describes one part of a life; the days where two of them meet are the
+// only place the record can say something neither source knows alone.
+type Overlap struct {
+ // Day is the calendar date.
+ Day time.Time
+ // Counts holds entries per source on that day.
+ Counts map[string]int
+ // Headlines holds one representative entry headline per source.
+ Headlines map[string]string
+}
+
+// Overlaps returns the days where every one of the given sources recorded
+// something, ordered by how much was recorded, heaviest first.
+func Overlaps(entries []vault.Entry, sources []string, now time.Time) []Overlap {
+ type bucket struct {
+ counts map[string]int
+ headlines map[string]string
+ day time.Time
+ }
+ byDay := map[string]*bucket{}
+ for _, e := range entries {
+ if !now.IsZero() && e.Time.After(now) {
+ continue
+ }
+ src := sourceOf(e, sources)
+ if src == "" {
+ continue
+ }
+ key := e.Time.Format("2006-01-02")
+ b := byDay[key]
+ if b == nil {
+ b = &bucket{
+ counts: map[string]int{},
+ headlines: map[string]string{},
+ day: time.Date(e.Time.Year(), e.Time.Month(), e.Time.Day(), 0, 0, 0, 0, e.Time.Location()),
+ }
+ byDay[key] = b
+ }
+ b.counts[src]++
+ if _, ok := b.headlines[src]; !ok {
+ b.headlines[src] = headline(e.Body)
+ }
+ }
+ var out []Overlap
+ for _, b := range byDay {
+ complete := true
+ for _, s := range sources {
+ if b.counts[s] == 0 {
+ complete = false
+ break
+ }
+ }
+ if !complete {
+ continue
+ }
+ out = append(out, Overlap{Day: b.day, Counts: b.counts, Headlines: b.headlines})
+ }
+ sort.Slice(out, func(i, j int) bool {
+ ti, tj := 0, 0
+ for _, n := range out[i].Counts {
+ ti += n
+ }
+ for _, n := range out[j].Counts {
+ tj += n
+ }
+ if ti != tj {
+ return ti > tj
+ }
+ return out[i].Day.Before(out[j].Day)
+ })
+ return out
+}
diff --git a/internal/weave/thread.go b/internal/weave/thread.go
new file mode 100644
index 0000000..ca31e7c
--- /dev/null
+++ b/internal/weave/thread.go
@@ -0,0 +1,451 @@
+// Package weave finds the shape of a record over time: what recurs, when it
+// started, when it stopped, and what took its place.
+//
+// The detection here is arithmetic rather than inference. A person can recall
+// what they did but cannot perceive absence, because nothing marks the last
+// time something happened. Counting occurrences and measuring silence against a
+// thread's own cadence surfaces those endings exactly, with no model in the
+// path and nothing to hallucinate.
+package weave
+
+import (
+ "regexp"
+ "sort"
+ "strings"
+ "time"
+
+ "github.com/dcadolph/midden/internal/vault"
+)
+
+// Status is where a thread stands at the observation date.
+type Status string
+
+// The statuses a thread can hold.
+const (
+ // Ongoing means the thread is still within its usual cadence.
+ Ongoing Status = "ongoing"
+ // Dormant means the thread is quiet but has already come back from a gap
+ // this long before, so the silence is its rhythm rather than its end.
+ Dormant Status = "dormant"
+ // Ended means the thread has been silent far longer than its cadence.
+ Ended Status = "ended"
+ // Emerging means the thread began recently and is still establishing.
+ Emerging Status = "emerging"
+)
+
+// Thread is a recurring pattern of entries sharing a normalized title.
+type Thread struct {
+ // Key is the normalized title threads are grouped by.
+ Key string
+ // Label is a representative original title, for display.
+ Label string
+ // Count is how many entries belong to the thread.
+ Count int
+ // First and Last are the earliest and latest occurrence.
+ First time.Time
+ Last time.Time
+ // MedianGap is the typical number of days between occurrences.
+ MedianGap int
+ // MaxGap is the longest the thread has ever gone quiet and come back. A
+ // yearly show and a weekly class both look silent in July; only this tells
+ // them apart.
+ MaxGap int
+ // SilentDays is how long the thread has gone quiet at the observation date.
+ SilentDays int
+ // SpanDays is how long the thread ran from first to last occurrence.
+ SpanDays int
+ // Status is where the thread stands.
+ Status Status
+ // Variants are the distinct headlines that were folded into this thread.
+ // A record whose whole worth is being true has to be auditable: a reader
+ // must be able to see what was grouped together before believing a claim
+ // built on the grouping.
+ Variants []string
+}
+
+// Weight ranks how much of a life a thread represents, combining how often it
+// happened with how long it ran. A weekly commitment held for two years outranks
+// a burst of activity over one month, which is the order a person would tell
+// them in.
+func (t Thread) Weight() int {
+ return t.Count * (t.SpanDays + 1)
+}
+
+// Options tune thread detection.
+type Options struct {
+ // Now is the observation date. Occurrences after it are ignored, so a
+ // calendar holding future events does not report them as current activity.
+ Now time.Time
+ // MinCount is the fewest occurrences a pattern needs to count as a thread.
+ MinCount int
+ // SilenceFactor multiplies a thread's own cadence to decide when silence
+ // means it ended. Measuring against the thread's own rhythm is what lets a
+ // daily habit and a yearly tradition be judged on the same scale.
+ SilenceFactor int
+ // MinSilenceDays floors the silence test, so a thread that happened twice in
+ // two days is not declared over by the third day.
+ MinSilenceDays int
+ // EmergingDays is how recently a thread must have started to count as new.
+ EmergingDays int
+ // DormantTolerance scales a thread's longest previous gap when deciding
+ // whether its current silence is seasonal rather than final.
+ DormantTolerance float64
+ // MergeSimilarity is how much two word sets must overlap to be treated as
+ // one thread, from zero to one. Word-set equality alone still splits a
+ // commitment recorded with an extra word attached, and each fragment then
+ // appears to end whenever the wording drifts.
+ MergeSimilarity float64
+}
+
+// DefaultOptions returns detection settings suited to a personal record.
+func DefaultOptions(now time.Time) Options {
+ return Options{
+ Now: now,
+ MinCount: 5,
+ SilenceFactor: 4,
+ MinSilenceDays: 90,
+ EmergingDays: 180,
+ MergeSimilarity: 0.6,
+ DormantTolerance: 1.3,
+ }
+}
+
+// Threads groups entries into recurring threads and classifies each one.
+// The result is ordered by weight, heaviest first.
+func Threads(entries []vault.Entry, opts Options) []Thread {
+ if opts.MinCount < 2 {
+ opts.MinCount = 2
+ }
+ if opts.SilenceFactor < 1 {
+ opts.SilenceFactor = 1
+ }
+ grouped := map[string][]time.Time{}
+ labels := map[string]string{}
+ variants := map[string]map[string]bool{}
+ for _, e := range entries {
+ if !opts.Now.IsZero() && e.Time.After(opts.Now) {
+ continue
+ }
+ key := NormalizeTitle(e.Body)
+ if key == "" {
+ continue
+ }
+ grouped[key] = append(grouped[key], e.Time)
+ if _, ok := labels[key]; !ok {
+ labels[key] = headline(e.Body)
+ }
+ if variants[key] == nil {
+ variants[key] = map[string]bool{}
+ }
+ variants[key][headline(e.Body)] = true
+ }
+ grouped, labels, variants = mergeSimilar(grouped, labels, variants, opts.MergeSimilarity)
+ out := make([]Thread, 0, len(grouped))
+ for key, times := range grouped {
+ if len(times) < opts.MinCount {
+ continue
+ }
+ t := buildThread(key, labels[key], times, opts)
+ t.Variants = sortedSet(variants[key])
+ out = append(out, t)
+ }
+ sort.Slice(out, func(i, j int) bool {
+ if out[i].Weight() != out[j].Weight() {
+ return out[i].Weight() > out[j].Weight()
+ }
+ return out[i].Key < out[j].Key
+ })
+ return out
+}
+
+// buildThread assembles one thread from its occurrence times.
+func buildThread(key, label string, times []time.Time, opts Options) Thread {
+ sort.Slice(times, func(i, j int) bool { return times[i].Before(times[j]) })
+ first, last := times[0], times[len(times)-1]
+ t := Thread{
+ Key: key,
+ Label: label,
+ Count: len(times),
+ First: first,
+ Last: last,
+ MedianGap: medianGapDays(times),
+ MaxGap: maxGapDays(times),
+ SpanDays: daysBetween(first, last),
+ }
+ if !opts.Now.IsZero() {
+ t.SilentDays = daysBetween(last, opts.Now)
+ }
+ t.Status = classify(t, opts)
+ return t
+}
+
+// classify decides where a thread stands. Silence is judged against the
+// thread's own cadence rather than a fixed window, so a weekly class and an
+// annual tradition are both measured fairly.
+func classify(t Thread, opts Options) Status {
+ limit := t.MedianGap * opts.SilenceFactor
+ if limit < opts.MinSilenceDays {
+ limit = opts.MinSilenceDays
+ }
+ if t.SilentDays > limit {
+ // A thread that has already returned from a gap this long is between
+ // seasons, not over. Calling that an ending would ask a person to
+ // explain the end of something that has not ended, which is worse than
+ // not asking: it asserts a false fact about their life.
+ if t.MaxGap > 0 && t.SilentDays <= int(float64(t.MaxGap)*opts.DormantTolerance) {
+ return Dormant
+ }
+ return Ended
+ }
+ if opts.EmergingDays > 0 && !opts.Now.IsZero() && daysBetween(t.First, opts.Now) <= opts.EmergingDays {
+ return Emerging
+ }
+ return Ongoing
+}
+
+// medianGapDays returns the typical number of days between sorted occurrences,
+// or zero when there are fewer than two.
+func medianGapDays(times []time.Time) int {
+ if len(times) < 2 {
+ return 0
+ }
+ gaps := make([]int, 0, len(times)-1)
+ for i := 1; i < len(times); i++ {
+ gaps = append(gaps, daysBetween(times[i-1], times[i]))
+ }
+ sort.Ints(gaps)
+ mid := len(gaps) / 2
+ if len(gaps)%2 == 1 {
+ return gaps[mid]
+ }
+ return (gaps[mid-1] + gaps[mid]) / 2
+}
+
+// maxGapDays returns the longest span between consecutive occurrences.
+func maxGapDays(times []time.Time) int {
+ longest := 0
+ for i := 1; i < len(times); i++ {
+ if g := daysBetween(times[i-1], times[i]); g > longest {
+ longest = g
+ }
+ }
+ return longest
+}
+
+// daysBetween returns whole days from a to b, never negative.
+func daysBetween(a, b time.Time) int {
+ d := int(b.Sub(a).Hours() / 24)
+ if d < 0 {
+ return 0
+ }
+ return d
+}
+
+// Title normalization. Calendar entries carry the same commitment written many
+// ways across years, so grouping has to survive case, emoji, punctuation, and
+// separator drift without merging genuinely different activities.
+var (
+ // parenSuffix strips the time or all-day marker the ingest appends.
+ parenSuffix = regexp.MustCompile(`\s*\((?:all day|\d{2}:\d{2} to \d{2}:\d{2})\)\s*$`)
+ // nonTitle drops anything that is not a letter, digit, space, or hyphen.
+ nonTitle = regexp.MustCompile(`[^\p{L}\p{N}\s-]+`)
+ // bareNumber drops standalone numbers such as a field or grade number.
+ bareNumber = regexp.MustCompile(`\b\d+\b`)
+ // spaces collapses runs of whitespace.
+ spaces = regexp.MustCompile(`\s+`)
+ // separator normalizes the dash or colon between a name and an activity.
+ separator = regexp.MustCompile(`\s*[-:]\s*`)
+)
+
+// filler holds the words that carry no identity in a calendar title. People
+// write the same commitment differently every time, so these are dropped before
+// grouping.
+var filler = map[string]bool{
+ "a": true, "an": true, "and": true, "at": true, "for": true, "from": true,
+ "in": true, "my": true, "of": true, "on": true, "our": true, "the": true,
+ "to": true, "w": true, "with": true, "his": true, "her": true, "their": true,
+ "s": true, "is": true, "be": true, "am": true, "pm": true,
+}
+
+// possessive strips a trailing possessive so a place named for a person matches
+// the person.
+var possessive = regexp.MustCompile(`(\p{L})['\x{2019}]s\b`)
+
+// NormalizeTitle reduces an entry body to the key its thread is grouped by.
+//
+// The key is a sorted set of meaningful words rather than the title itself,
+// because a handwritten calendar records one commitment under many spellings.
+// "sleepover with Kayla", "sleepover @ Kayla's", and "sleepover w/ Kayla" are
+// the same standing arrangement, and grouping on the exact string splits them
+// into separate threads that each appear to end the moment the wording changes.
+// Collapsing to a word set makes the phrasing irrelevant and the subject decisive.
+func NormalizeTitle(body string) string {
+ t := strings.ToLower(headline(body))
+ t = parenSuffix.ReplaceAllString(t, "")
+ t = possessive.ReplaceAllString(t, "$1")
+ t = nonTitle.ReplaceAllString(t, " ")
+ t = separator.ReplaceAllString(t, " ")
+ t = bareNumber.ReplaceAllString(t, " ")
+ t = spaces.ReplaceAllString(t, " ")
+
+ seen := map[string]bool{}
+ var words []string
+ for _, w := range strings.Fields(t) {
+ w = strings.Trim(w, "-")
+ if len(w) < 2 || filler[w] || seen[w] {
+ continue
+ }
+ seen[w] = true
+ words = append(words, w)
+ }
+ if len(words) == 0 {
+ return ""
+ }
+ sort.Strings(words)
+ return strings.Join(words, " ")
+}
+
+// headline returns the first non-empty line of an entry body.
+func headline(body string) string {
+ for line := range strings.SplitSeq(body, "\n") {
+ if s := strings.TrimSpace(line); s != "" {
+ return parenSuffix.ReplaceAllString(s, "")
+ }
+ }
+ return ""
+}
+
+// mergeSimilar folds together groups whose word sets overlap enough to be the
+// same thread. Occurrences are merged into the largest group, which also keeps
+// the most representative label.
+func mergeSimilar(
+ grouped map[string][]time.Time,
+ labels map[string]string,
+ variants map[string]map[string]bool,
+ threshold float64,
+) (map[string][]time.Time, map[string]string, map[string]map[string]bool) {
+ if threshold <= 0 || threshold > 1 {
+ return grouped, labels, variants
+ }
+ keys := make([]string, 0, len(grouped))
+ for k := range grouped {
+ keys = append(keys, k)
+ }
+ // Largest first, so smaller variants fold into the dominant spelling.
+ sort.Slice(keys, func(i, j int) bool {
+ if len(grouped[keys[i]]) != len(grouped[keys[j]]) {
+ return len(grouped[keys[i]]) > len(grouped[keys[j]])
+ }
+ return keys[i] < keys[j]
+ })
+ sets := make([]map[string]bool, len(keys))
+ for i, k := range keys {
+ sets[i] = wordSet(k)
+ }
+ outTimes := map[string][]time.Time{}
+ outLabels := map[string]string{}
+ outVariants := map[string]map[string]bool{}
+ canonical := make([]int, 0, len(keys))
+ for i, k := range keys {
+ target := -1
+ for _, c := range canonical {
+ if jaccard(sets[i], sets[c]) >= threshold && !addsSubject(sets[i], sets[c]) {
+ target = c
+ break
+ }
+ }
+ if target < 0 {
+ canonical = append(canonical, i)
+ outTimes[k] = append(outTimes[k], grouped[k]...)
+ outLabels[k] = labels[k]
+ outVariants[k] = copySet(variants[k])
+ continue
+ }
+ ck := keys[target]
+ outTimes[ck] = append(outTimes[ck], grouped[k]...)
+ if outVariants[ck] == nil {
+ outVariants[ck] = map[string]bool{}
+ }
+ for v := range variants[k] {
+ outVariants[ck][v] = true
+ }
+ }
+ return outTimes, outLabels, outVariants
+}
+
+// copySet duplicates a set so later merges cannot mutate the original.
+func copySet(in map[string]bool) map[string]bool {
+ out := make(map[string]bool, len(in))
+ for k := range in {
+ out[k] = true
+ }
+ return out
+}
+
+// sortedSet renders a set as a stable ordered slice.
+func sortedSet(in map[string]bool) []string {
+ out := make([]string, 0, len(in))
+ for k := range in {
+ out = append(out, k)
+ }
+ sort.Strings(out)
+ return out
+}
+
+// addsSubject reports whether merging two word sets would absorb a title that
+// names nobody into one that names someone.
+//
+// "Doctor appt" and "Hannah doctor appt" overlap heavily, but the second says
+// whose appointment it was and the first does not. Folding them together treats
+// two people's appointments as one thread, which then reports a gap spanning the
+// distance between two unrelated lives. A short title has no room for a spare
+// word, so a word added to one is a subject rather than noise; a longer title
+// can absorb one without changing what it is about.
+func addsSubject(a, b map[string]bool) bool {
+ shorter, longer := a, b
+ if len(b) < len(a) {
+ shorter, longer = b, a
+ }
+ if len(shorter) >= minTokensForNoise {
+ return false
+ }
+ for w := range shorter {
+ if !longer[w] {
+ // Not a subset: the two differ in both directions, so neither is a
+ // bare version of the other.
+ return false
+ }
+ }
+ return len(longer) > len(shorter)
+}
+
+// minTokensForNoise is how many words a title needs before an extra one can be
+// treated as incidental rather than as the subject.
+const minTokensForNoise = 3
+
+// wordSet splits a normalized key back into its words.
+func wordSet(key string) map[string]bool {
+ out := map[string]bool{}
+ for _, w := range strings.Fields(key) {
+ out[w] = true
+ }
+ return out
+}
+
+// jaccard returns the overlap of two word sets as intersection over union.
+func jaccard(a, b map[string]bool) float64 {
+ if len(a) == 0 || len(b) == 0 {
+ return 0
+ }
+ inter := 0
+ for w := range a {
+ if b[w] {
+ inter++
+ }
+ }
+ union := len(a) + len(b) - inter
+ if union == 0 {
+ return 0
+ }
+ return float64(inter) / float64(union)
+}
diff --git a/internal/weave/thread_test.go b/internal/weave/thread_test.go
new file mode 100644
index 0000000..e597cd0
--- /dev/null
+++ b/internal/weave/thread_test.go
@@ -0,0 +1,378 @@
+package weave
+
+import (
+ "fmt"
+ "testing"
+ "time"
+
+ "github.com/google/go-cmp/cmp"
+ "github.com/google/go-cmp/cmp/cmpopts"
+
+ "github.com/dcadolph/midden/internal/vault"
+)
+
+// on builds an entry with the given body on the given date.
+func on(year int, month time.Month, day int, body string) vault.Entry {
+ return vault.Entry{Time: time.Date(year, month, day, 12, 0, 0, 0, time.Local), Body: body}
+}
+
+// weekly builds n entries seven days apart starting at the given date.
+func weekly(year int, month time.Month, day, n int, body string) []vault.Entry {
+ out := make([]vault.Entry, 0, n)
+ start := time.Date(year, month, day, 12, 0, 0, 0, time.Local)
+ for i := range n {
+ d := start.AddDate(0, 0, 7*i)
+ out = append(out, vault.Entry{Time: d, Body: body})
+ }
+ return out
+}
+
+// observed is a fixed observation date, so classification never depends on when
+// the suite runs.
+var observed = time.Date(2026, time.August, 26, 12, 0, 0, 0, time.Local)
+
+// find returns the thread whose label matches, or fails the test.
+func find(t *testing.T, threads []Thread, label string) Thread {
+ t.Helper()
+ for _, th := range threads {
+ if th.Label == label {
+ return th
+ }
+ }
+ t.Fatalf("no thread labeled %q in %d threads", label, len(threads))
+ return Thread{}
+}
+
+func TestNormalizeTitleMergesPhrasingVariants(t *testing.T) {
+ t.Parallel()
+ tests := []struct {
+ Name string
+ In []string
+ Same bool
+ }{{ // Test 0: A handwritten commitment survives separator and preposition drift.
+ Name: "sleepover variants",
+ In: []string{
+ "Hannah- sleepover with Kayla",
+ "Hannah- sleepover @ Kayla's",
+ "Hannah - sleepover at Kayla’s (18:00 to 13:00)",
+ "Hannah- sleepover w/ Kayla",
+ },
+ Same: true,
+ }, { // Test 1: Case and emoji do not split a thread.
+ Name: "case and emoji",
+ In: []string{"William- Martial arts", "William- martial arts ⚔️", "WILLIAM - MARTIAL ARTS"},
+ Same: true,
+ }, { // Test 2: A trailing field or grade number does not split a thread.
+ Name: "trailing number",
+ In: []string{"William- baseball practice (field 3)", "William- baseball practice field 4"},
+ Same: true,
+ }, { // Test 3: Genuinely different activities stay apart.
+ Name: "distinct activities",
+ In: []string{"Hannah- hip hop", "Hannah- acro"},
+ Same: false,
+ }, { // Test 4: The same activity for different people stays apart.
+ Name: "different subjects",
+ In: []string{"William- piano", "Haylie- piano"},
+ Same: false,
+ }}
+ for testNum, test := range tests {
+ t.Run(fmt.Sprintf("test %d %s", testNum, test.Name), func(t *testing.T) {
+ t.Parallel()
+ first := NormalizeTitle(test.In[0])
+ for _, other := range test.In[1:] {
+ got := NormalizeTitle(other)
+ if test.Same && got != first {
+ t.Errorf("want %q and %q to share a key, got %q vs %q", test.In[0], other, first, got)
+ }
+ if !test.Same && got == first {
+ t.Errorf("want %q and %q to stay separate, both keyed %q", test.In[0], other, got)
+ }
+ }
+ })
+ }
+}
+
+func TestThreadsClassifiesByOwnCadence(t *testing.T) {
+ t.Parallel()
+ var entries []vault.Entry
+ // A weekly class that stopped a year ago.
+ entries = append(entries, weekly(2024, time.January, 6, 40, "William- martial arts")...)
+ // A weekly class still running up to the observation date.
+ entries = append(entries, weekly(2026, time.June, 3, 12, "William- baseball practice")...)
+ // A yearly tradition, which must not read as ended merely because a year passed.
+ for y := 2022; y <= 2026; y++ {
+ entries = append(entries, on(y, time.March, 4, "Dad birthday"))
+ }
+ opts := DefaultOptions(observed)
+ threads := Threads(entries, opts)
+
+ if got := find(t, threads, "William- martial arts"); got.Status != Ended {
+ t.Errorf("want a long-silent weekly class to read as ended, got %s (silent %d days)", got.Status, got.SilentDays)
+ }
+ if got := find(t, threads, "William- baseball practice"); got.Status == Ended {
+ t.Errorf("want a current weekly class to read as live, got %s", got.Status)
+ }
+ // A yearly gap is normal for a yearly thread, which is the whole reason
+ // silence is measured against cadence rather than a fixed window.
+ if got := find(t, threads, "Dad birthday"); got.Status == Ended {
+ t.Errorf("want a yearly tradition to survive a yearly gap, got ended (gap %d days)", got.MedianGap)
+ }
+}
+
+func TestThreadsIgnoresFutureEntries(t *testing.T) {
+ t.Parallel()
+ var entries []vault.Entry
+ entries = append(entries, weekly(2026, time.August, 5, 3, "Dentist")...)
+ // Calendars hold future appointments; counting them would report activity
+ // that has not happened.
+ entries = append(entries, on(2027, time.March, 12, "Dentist"))
+ threads := Threads(entries, Options{Now: observed, MinCount: 2, SilenceFactor: 4, MinSilenceDays: 90})
+ got := find(t, threads, "Dentist")
+ if got.Count != 3 {
+ t.Errorf("want 3 past occurrences, got %d", got.Count)
+ }
+ if got.Last.After(observed) {
+ t.Errorf("want the last occurrence on or before the observation date, got %s", got.Last)
+ }
+}
+
+func TestThreadsHonorsMinCount(t *testing.T) {
+ t.Parallel()
+ entries := append(weekly(2026, time.July, 1, 6, "Recurring"), on(2026, time.July, 2, "One off"))
+ threads := Threads(entries, DefaultOptions(observed))
+ for _, th := range threads {
+ if th.Label == "One off" {
+ t.Error("want a single occurrence excluded from threads")
+ }
+ }
+ find(t, threads, "Recurring")
+}
+
+func TestMergeSimilarFoldsExtraWords(t *testing.T) {
+ t.Parallel()
+ // The same standing arrangement, recorded once with an extra name attached.
+ entries := []vault.Entry{
+ on(2025, time.January, 4, "Hannah- sleepover with Kayla"),
+ on(2025, time.February, 8, "Hannah- sleepover @ Kayla's"),
+ on(2025, time.March, 15, "Hannah- sleepover w/ Kayla"),
+ on(2025, time.April, 19, "Hannah- sleepover with Kayla"),
+ on(2025, time.May, 24, "Hannah- Kayla sleepover"),
+ }
+ threads := Threads(entries, Options{
+ Now: observed, MinCount: 5, SilenceFactor: 4, MinSilenceDays: 90, MergeSimilarity: 0.6,
+ })
+ if len(threads) != 1 {
+ t.Fatalf("want one merged thread, got %d: %v", len(threads), labelsOf(threads))
+ }
+ if threads[0].Count != 5 {
+ t.Errorf("want all 5 occurrences merged, got %d", threads[0].Count)
+ }
+}
+
+func TestMergeSimilarKeepsDistinctThreadsApart(t *testing.T) {
+ t.Parallel()
+ var entries []vault.Entry
+ entries = append(entries, weekly(2026, time.January, 6, 6, "Hannah- hip hop")...)
+ entries = append(entries, weekly(2026, time.January, 7, 6, "William- baseball practice")...)
+ threads := Threads(entries, DefaultOptions(observed))
+ if len(threads) != 2 {
+ t.Errorf("want two distinct threads, got %d: %v", len(threads), labelsOf(threads))
+ }
+}
+
+func TestMedianGapDays(t *testing.T) {
+ t.Parallel()
+ base := time.Date(2026, time.January, 1, 0, 0, 0, 0, time.Local)
+ times := []time.Time{base, base.AddDate(0, 0, 7), base.AddDate(0, 0, 14), base.AddDate(0, 0, 21)}
+ if diff := cmp.Diff(7, medianGapDays(times)); diff != "" {
+ t.Errorf("mismatch (-want +got):\n%s", diff)
+ }
+ if diff := cmp.Diff(0, medianGapDays(times[:1])); diff != "" {
+ t.Errorf("single occurrence should have no gap (-want +got):\n%s", diff)
+ }
+}
+
+func TestWeightPrefersLongRunningThreads(t *testing.T) {
+ t.Parallel()
+ long := Thread{Count: 50, SpanDays: 700}
+ burst := Thread{Count: 50, SpanDays: 20}
+ // Two threads of equal volume are not equal in a life: the one held for two
+ // years mattered more than the one that filled three weeks.
+ if long.Weight() <= burst.Weight() {
+ t.Errorf("want the sustained thread to outrank the burst, got %d vs %d", long.Weight(), burst.Weight())
+ }
+}
+
+func TestThreadsEmptyInput(t *testing.T) {
+ t.Parallel()
+ if diff := cmp.Diff([]Thread(nil), Threads(nil, DefaultOptions(observed)), cmpopts.EquateEmpty()); diff != "" {
+ t.Errorf("mismatch (-want +got):\n%s", diff)
+ }
+}
+
+// labelsOf returns thread labels for failure messages.
+func labelsOf(threads []Thread) []string {
+ out := make([]string, len(threads))
+ for i, t := range threads {
+ out[i] = t.Label
+ }
+ return out
+}
+
+func TestClassifyDormantForSeasonalThreads(t *testing.T) {
+ t.Parallel()
+ var entries []vault.Entry
+ // A show that happens each spring. In August it is silent by months, but it
+ // has returned from a gap this long every year of its life.
+ for y := 2023; y <= 2026; y++ {
+ entries = append(entries, on(y, time.April, 25, "Haylie- spring show"))
+ entries = append(entries, on(y, time.April, 26, "Haylie- spring show"))
+ }
+ threads := Threads(entries, Options{
+ Now: observed, MinCount: 5, SilenceFactor: 4, MinSilenceDays: 90,
+ DormantTolerance: 1.3, MergeSimilarity: 0.6,
+ })
+ got := find(t, threads, "Haylie- spring show")
+ if got.Status != Dormant {
+ t.Errorf("want a seasonal thread read as dormant, got %s (silent %d, max gap %d)",
+ got.Status, got.SilentDays, got.MaxGap)
+ }
+}
+
+func TestClassifyEndedWhenSilenceExceedsAnyPreviousGap(t *testing.T) {
+ t.Parallel()
+ // A weekly class that ran steadily and then stopped for far longer than it
+ // ever paused. Nothing about its history explains this silence.
+ entries := weekly(2024, time.January, 6, 40, "William- martial arts")
+ threads := Threads(entries, DefaultOptions(observed))
+ got := find(t, threads, "William- martial arts")
+ if got.Status != Ended {
+ t.Errorf("want a genuinely stopped thread read as ended, got %s (silent %d, max gap %d)",
+ got.Status, got.SilentDays, got.MaxGap)
+ }
+}
+
+func TestMaxGapRecordsTheLongestReturn(t *testing.T) {
+ t.Parallel()
+ entries := []vault.Entry{
+ on(2024, time.January, 1, "Thing"),
+ on(2024, time.January, 8, "Thing"),
+ // A long pause, then it comes back.
+ on(2025, time.January, 8, "Thing"),
+ on(2025, time.January, 15, "Thing"),
+ on(2025, time.January, 22, "Thing"),
+ }
+ threads := Threads(entries, Options{Now: observed, MinCount: 5, SilenceFactor: 4, MinSilenceDays: 90})
+ got := find(t, threads, "Thing")
+ if got.MaxGap < 360 {
+ t.Errorf("want the year-long pause recorded, got %d days", got.MaxGap)
+ }
+ if got.MedianGap > 30 {
+ t.Errorf("want the median to stay near the usual weekly rhythm, got %d days", got.MedianGap)
+ }
+}
+
+func TestMergeDoesNotAbsorbATitleThatNamesNobody(t *testing.T) {
+ t.Parallel()
+ var entries []vault.Entry
+ // One person's appointments, recorded several ways.
+ for i := range 6 {
+ entries = append(entries, on(2024, time.Month(1+i), 10, "Hannah- doctor appt"))
+ }
+ // A generic appointment naming nobody, years earlier. It could be anyone's.
+ for i := range 6 {
+ entries = append(entries, on(2015, time.Month(1+i), 10, "Doctor appt"))
+ }
+ threads := Threads(entries, DefaultOptions(observed))
+ // Folding these together would treat two people's appointments as one
+ // thread and report a gap spanning the distance between two unrelated lives.
+ for _, th := range threads {
+ if len(th.Variants) > 1 {
+ t.Errorf("want the unnamed appointments kept separate, merged: %v", th.Variants)
+ }
+ }
+ if len(threads) != 2 {
+ t.Errorf("want two distinct threads, got %d: %v", len(threads), labelsOf(threads))
+ }
+}
+
+func TestMergeStillFoldsPhrasingOfOneCommitment(t *testing.T) {
+ t.Parallel()
+ // Every one of these names the same two people doing the same thing, so an
+ // extra word is drift rather than a change of subject.
+ entries := []vault.Entry{
+ on(2025, time.January, 4, "Hannah- sleepover with Kayla"),
+ on(2025, time.February, 8, "Hannah- sleepover @ Kayla's"),
+ on(2025, time.March, 15, "Hannah- sleepover w/ Kayla"),
+ on(2025, time.April, 19, "Hannah- Kayla and Sara sleepover"),
+ on(2025, time.May, 24, "Hannah - sleepover with Kayla"),
+ }
+ threads := Threads(entries, DefaultOptions(observed))
+ if len(threads) != 1 {
+ t.Fatalf("want one thread, got %d: %v", len(threads), labelsOf(threads))
+ }
+ if threads[0].Count != 5 {
+ t.Errorf("want all five occurrences merged, got %d", threads[0].Count)
+ }
+ if len(threads[0].Variants) != 5 {
+ t.Errorf("want every phrasing recorded for audit, got %v", threads[0].Variants)
+ }
+}
+
+func TestVariantsAreRecordedAndSorted(t *testing.T) {
+ t.Parallel()
+ entries := []vault.Entry{
+ on(2025, time.January, 4, "William- martial arts"),
+ on(2025, time.January, 11, "William- Martial arts"),
+ on(2025, time.January, 18, "William- martial arts ⚔️"),
+ on(2025, time.January, 25, "William- martial arts"),
+ on(2025, time.February, 1, "William- martial arts"),
+ }
+ threads := Threads(entries, DefaultOptions(observed))
+ if len(threads) != 1 {
+ t.Fatalf("want one thread, got %d", len(threads))
+ }
+ got := threads[0].Variants
+ if len(got) != 3 {
+ t.Errorf("want the three distinct spellings recorded, got %v", got)
+ }
+ for i := 1; i < len(got); i++ {
+ if got[i] < got[i-1] {
+ t.Errorf("want variants in a stable order, got %v", got)
+ }
+ }
+}
+
+func TestAddsSubject(t *testing.T) {
+ t.Parallel()
+ set := func(words ...string) map[string]bool {
+ m := map[string]bool{}
+ for _, w := range words {
+ m[w] = true
+ }
+ return m
+ }
+ tests := []struct {
+ Name string
+ A map[string]bool
+ B map[string]bool
+ Want bool
+ }{
+ {"short title gains a name", set("appt", "doctor"), set("appt", "doctor", "hannah"), true},
+ {"longer title gains a word", set("hannah", "kayla", "sleepover"), set("hannah", "kayla", "sleepover", "sara"), false},
+ {"neither is a subset", set("hannah", "acro"), set("hannah", "piano"), false},
+ {"identical sets", set("hannah", "acro"), set("hannah", "acro"), false},
+ }
+ for _, test := range tests {
+ t.Run(test.Name, func(t *testing.T) {
+ t.Parallel()
+ if got := addsSubject(test.A, test.B); got != test.Want {
+ t.Errorf("want %v, got %v", test.Want, got)
+ }
+ // The check must not depend on argument order.
+ if got := addsSubject(test.B, test.A); got != test.Want {
+ t.Errorf("want %v regardless of order, got %v", test.Want, got)
+ }
+ })
+ }
+}
diff --git a/llm/chat.go b/llm/chat.go
index 56ccbb3..1bd588f 100644
--- a/llm/chat.go
+++ b/llm/chat.go
@@ -116,8 +116,11 @@ func (c *anthropicChat) Reply(ctx context.Context, system string, history []Mess
return "", err
}
var parsed struct {
- StopReason string `json:"stop_reason"`
- Content []struct {
+ StopReason string `json:"stop_reason"`
+ StopDetails struct {
+ Category string `json:"category"`
+ } `json:"stop_details"`
+ Content []struct {
Type string `json:"type"`
Text string `json:"text"`
} `json:"content"`
@@ -125,6 +128,15 @@ func (c *anthropicChat) Reply(ctx context.Context, system string, history []Mess
if err := json.Unmarshal(data, &parsed); err != nil {
return "", fmt.Errorf("decode response: %w", err)
}
+ // A declined request is a successful HTTP response carrying no content, so
+ // reading the blocks without checking would return an empty reply as if the
+ // journal simply had nothing to say.
+ if parsed.StopReason == "refusal" {
+ if parsed.StopDetails.Category != "" {
+ return "", fmt.Errorf("anthropic declined the request (%s)", parsed.StopDetails.Category)
+ }
+ return "", errors.New("anthropic declined the request")
+ }
var b strings.Builder
for _, p := range parsed.Content {
if p.Type == "text" {
diff --git a/llm/client.go b/llm/client.go
index 68de24e..d6be79a 100644
--- a/llm/client.go
+++ b/llm/client.go
@@ -18,16 +18,19 @@ import (
const EnvChatMaxTokens = "MIDDEN_CHAT_MAX_TOKENS" //nolint:gosec // Environment variable name, not a credential.
// Default provider models and endpoints. They live in one block so provider
-// drift is a one-line fix; each has an environment override.
+// drift is a one-line fix; each has an environment override. The reply budget
+// is generous because current Claude models reason before answering and that
+// reasoning is charged against the same cap as the reply, so a tight budget
+// truncates the answer rather than the thinking.
const (
- defaultAnthropicModel = "claude-sonnet-4-6"
+ defaultAnthropicModel = "claude-opus-5"
defaultOpenAIChatModel = "gpt-4o-mini"
defaultOllamaChatModel = "llama3.2"
defaultOpenAIEmbedModel = "text-embedding-3-small"
defaultVoyageEmbedModel = "voyage-3"
defaultOllamaEmbedModel = "nomic-embed-text"
defaultOllamaHost = "http://localhost:11434"
- defaultChatMaxTokens = 4096
+ defaultChatMaxTokens = 16000
)
// retryAttempts is the total try count for retryable provider failures.
diff --git a/llm/llm_test.go b/llm/llm_test.go
index dd65ed6..33342df 100644
--- a/llm/llm_test.go
+++ b/llm/llm_test.go
@@ -184,3 +184,35 @@ func TestChatMaxTokensOverride(t *testing.T) {
t.Errorf("invalid value must fall back to default, got %d", got)
}
}
+
+func TestAnthropicChatSurfacesRefusal(t *testing.T) {
+ stubHTTP(t, stubResponse{
+ Status: 200,
+ Body: `{"stop_reason":"refusal","stop_details":{"type":"refusal","category":"cyber"},"content":[]}`,
+ })
+ c := &anthropicChat{apiKey: "test", model: "claude-opus-5", maxTokens: 1024}
+ // A refusal is an HTTP 200 with no content blocks, so reading the blocks
+ // without checking would report an empty answer as a real one.
+ reply, err := c.Reply(context.Background(), "system", []Message{{Role: "user", Content: "hi"}})
+ if err == nil {
+ t.Fatalf("want an error for a refused request, got reply %q", reply)
+ }
+ if !strings.Contains(err.Error(), "cyber") {
+ t.Errorf("want the refusal category surfaced, got %v", err)
+ }
+}
+
+func TestAnthropicChatReturnsText(t *testing.T) {
+ stubHTTP(t, stubResponse{
+ Status: 200,
+ Body: `{"stop_reason":"end_turn","content":[{"type":"thinking","text":""},{"type":"text","text":"answer"}]}`,
+ })
+ c := &anthropicChat{apiKey: "test", model: "claude-opus-5", maxTokens: 1024}
+ reply, err := c.Reply(context.Background(), "system", []Message{{Role: "user", Content: "hi"}})
+ if err != nil {
+ t.Fatalf("Reply: %v", err)
+ }
+ if diff := cmp.Diff("answer", reply); diff != "" {
+ t.Errorf("mismatch (-want +got):\n%s", diff)
+ }
+}
diff --git a/skill/SKILL.md b/skill/SKILL.md
index 5df108a..8ae43a7 100644
--- a/skill/SKILL.md
+++ b/skill/SKILL.md
@@ -8,7 +8,8 @@ description: >
"remember X for later", "note that", "jot down", "don't forget", "midden add",
"/journal", "what did I write about X", "when did I last", "show me recent
entries", "what did I do on ", "flashback", "on this day", "how is my
- streak", "search my journal", "list tags".
+ streak", "search my journal", "list tags", "what happened last ",
+ "what do you know about my life", "import my calendar", "backfill my history".
---
Midden is a personal markdown journal on the user's machine. Daily files live at
@@ -57,7 +58,7 @@ Map the user's question to the smallest matching command:
| User wants | Command |
|---|---|
-| Last single entry | `midden last` |
+| Last single entry | `midden last` (stops at now; `--future` for scheduled entries) |
| Last N entries | `midden last -n N` or `midden recent -n N` |
| Everything on a date | `midden on YYYY-MM-DD` (or `today`, `yesterday`, weekday names, `N-units-ago`) |
| Everything in a range | `midden between FROM TO` |
@@ -77,10 +78,21 @@ Map the user's question to the smallest matching command:
| Import a markdown file as an entry | `midden import path/to/file.md --tag inbox` (or `-` to read from stdin) |
| Render an HTML report | `midden report html -o ~/midden-report.html` |
| Semantic search ("anything about token rotation") | `midden recall "token rotation"` (requires a prior `midden reindex`) |
-| Synthesize an answer from journal evidence | `midden chat "when did I last see Mom"` |
-| Rebuild the embedding index | `midden reindex` |
+| Synthesize an answer about a topic | `midden chat "when did I last see Mom"` |
+| Answer a question about a period | `midden chat --since 2026-03-01 --until 2026-03-31 "what happened"` |
+| Answer a question about the whole record | `midden chat --sweep "what do you know about my life"` (add `--context-chars 6000` for a small local model) |
+| What recurs, started, or stopped | `midden weave --tag calendar` (threads, handoffs, crossings) |
+| Prompt the user about a gap in their record | `midden ask` to see it, `midden ask --answer "..."` to record a reply |
+| Rebuild the embedding index | `midden reindex` (reuses unchanged vectors; `--full` re-embeds everything) |
| Capture a voice memo (optionally transcribed) | `midden audio --duration 30s --transcribe` |
-| Ingest an .ics calendar export | `midden ingest ics ~/Downloads/cal.ics --from today --to today` |
+| Ingest an .ics calendar export | `midden ingest ics ~/Downloads/cal.ics` (whole file; narrow with `--since`/`--until`) |
+| Ingest commit history from repositories | `midden ingest git ~/src/project` (add `--author`, `--since`, `--stat`) |
+
+A vault holding an imported calendar contains appointments that have not happened
+yet. `last` and `recent` stop at the present so they do not report next spring's
+dentist appointment as the most recent thing in the record; pass `--future` when
+the user is asking what is coming up. `stats` reports scheduled entries on their
+own line for the same reason.
Pass `--json` to any of the read commands to receive structured output you can
parse without regex.
@@ -107,15 +119,105 @@ argument list is visible to other processes.
## LLM-backed recall and chat
`midden recall ""` performs a semantic search via the embedding index;
-`midden chat ""` adds an LLM synthesis step that quotes journal entries
-as evidence. Both require a prior `midden reindex` and at least one provider
+`midden chat ""` adds a synthesis step that quotes journal entries as
+evidence. Both require a prior `midden reindex` and at least one provider
configured through environment variables (`VOYAGE_API_KEY`, `OPENAI_API_KEY`,
-`ANTHROPIC_API_KEY`, or local Ollama).
+`ANTHROPIC_API_KEY`, or local Ollama). Embedding and chat providers are chosen
+independently via `MIDDEN_EMBED_PROVIDER` and `MIDDEN_CHAT_PROVIDER`, so a local
+embedder can be paired with a hosted chat model.
+
+### Choose the retrieval shape before running chat
+
+This matters more than which command you pick. Ranking by similarity answers
+"what did I write about X" well and "what happened last March" badly, because the
+entries closest to a question about a period are often from other periods.
+
+- **Topic question** ("anything about the token rotation work"): plain
+ `midden chat "..."`. Similarity ranking is the right tool.
+- **Period question** ("what happened last March", "what did I do this month",
+ "how was that trip"): always pass `--since` and `--until`. Scoping switches
+ chat from reading the closest few entries to reading every entry in the range,
+ summarizing in chunks when the range is large. Without it the answer is drawn
+ from the wrong dates and will look plausible while being wrong.
+- **Whole-record question** ("what do you know about my life", "what people and
+ places seem important", "what do I keep coming back to"): pass `--sweep`.
+ Nearest-neighbor search cannot answer a question about the shape of a record.
+
+Both flags accept any date midden understands: `2026-03-01`, `30-days-ago`,
+`monday`, `today`.
+
+A sweep over a large range reads everything in it, summarizing in chunks when the
+range exceeds what the model can read at once. `--context-chars` sets that size.
+The default suits a hosted model. When the configured chat provider is a small
+local model, pass a much lower value (around `6000`) or the sweep will appear to
+hang; midden prints per-chunk progress while it works.
+
+Every chat answer also receives a summary counted over the entire indexed record
+in scope: how many entries, the span they cover, the tag histogram, and entries
+per month. Those counts are complete even when the quoted entries are a sample,
+so trust them for questions about frequency and shape, and never conclude
+something did not happen merely because it is missing from the quoted entries.
Prefer `midden recall` when the user wants to find entries.
Prefer `midden chat` when the user wants a narrated answer that cites entries.
If recall returns nothing useful, fall back to `midden search` over the raw text.
+### Weave: what the user cannot ask for
+
+Use `midden weave` when the user asks what has changed, what they have stopped
+doing, what is new, or asks an open question about their own life over time.
+Recall and chat can only surface what the user already knows to ask about;
+weave reports things nobody wrote down, chiefly endings, because nothing marks
+the last time something happened.
+
+Pass `--tag calendar` (or whichever life source the vault holds) when the vault
+also contains commit history, or repeated commit subjects will crowd out the
+real threads. Read the sections as: `Silences` are stretches where the whole record
+went quiet and are the largest thing it can be missing, `Ended` is what went quiet, `Started` is
+what is new, `Ongoing` is the steady weight of the record, `Handoffs` are
+successions where one thread stopped and another began, and `Crossings` are days
+where two sources meet.
+
+Every number weave prints is counted, not inferred, so quote them exactly and do
+not embellish. Use `midden weave --explain` to show which headlines were folded
+into a thread when a grouping looks doubtful. Do check a surprising ending before presenting it as fact: run
+`midden search` on the subject to confirm the thread really stopped rather than
+being recorded under different wording.
+
+### Ask: capturing what no import can reach
+
+Imported history covers where the user was and what they produced. It never
+covers what they thought, and for anyone who did not already journal there is
+nothing to import that would. `midden ask` is how that gap closes.
+
+Run `midden ask` when the user asks what they should record, wants a prompt, or
+finishes a backfill and wonders what to do next. Put the question and its
+evidence to them verbatim; the evidence is what makes the question answerable.
+Record the reply with `midden ask --answer ""`, using their words
+rather than a summary, since the point of the entry is their voice. Use
+`midden ask --skip` when the user says a question is not worth answering, so it
+leaves the queue for good rather than being put to them again.
+
+Never invent an answer, and never file a plausible-sounding reply on the user's
+behalf. A fabricated entry is worse than a missing one: the whole value of the
+record is that everything in it is true, and a counterfeit memory is
+indistinguishable from a real one once it is written.
+
+### Backfilling history
+
+When a user wants to populate the vault from records they already have, prefer
+ingestion over asking them to write anything:
+
+- `midden ingest ics ` imports a calendar export. Recurring series are
+ expanded into the occurrences they actually produced, cancellations are
+ honored, and rescheduled instances replace the occurrence they override.
+ Re-running it changes nothing, so it is safe to repeat.
+- `midden ingest git ...` imports commit history across any number of
+ repositories, filed by author date and deduplicated by hash. Use `--author` to
+ narrow a shared repository to the user's own commits.
+
+Run `midden reindex` after any ingest so recall and chat can see the new entries.
+
## Backups
When the user asks to back up, sync, or version their vault, prefer