Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
14 changes: 14 additions & 0 deletions TESTING_GUIDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -56,6 +56,13 @@ Also cover index-ordered chunk metadata, bounded legacy fallback, offset-only fi
page authority, and combined execution/output caps without synthesized continuation
in the CLI and both MCP tools.

`FindChunkReadTests` (#5415) shares a many-chunk fixture across literal, exact,
trigram-candidate and line-regex row/count queries on current and read-only legacy
indexes. Keep ordered query-plan checks, bounded metadata pages, reversed insertion,
overlap ownership, missing/live source controls, NULL-content failures, line caps,
and row/count continuation across chunk pages. Run on net8/net9 alongside existing
find, pagination and MCP regressions; no peak-memory or latency claim is implied.

`InspectCompactCandidateTests` (#5397) shares a real indexed fixture across CLI
inspect and MCP compact analysis. Keep zero, exact-limit, probe-overflow and the
internal five-definition boundary, partial and unrelated same-name/overload
Expand Down Expand Up @@ -1867,6 +1874,13 @@ UTF-16 座標、重複チャンク、一致範囲の上限内外、アンカー
チャンクの索引順取得と旧索引の上限付き代替処理、offset だけで再開した最終ページの確定性、
実行上限と出力上限が同時に発生しても継続カーソルを作らない規則を CLI と両 MCP ツールで検証します。

`FindChunkReadTests`(#5415)は多数のチャンクを持つ共通 fixture を使い、現在の索引と
読取専用の旧索引でリテラル・exact・trigram 候補・行単位正規表現の行/件数検索を検証します。
順序付きクエリ計画、上限付きメタデータページ、逆順挿入、重複行の優先順、索引本文の欠落と
実ファイルの対照、NULL 本文エラー、行上限、チャンクページをまたぐ行/件数の継続取得を
維持してください。既存の find・ページ分割・MCP 回帰とともに net8/net9 で実行します。
ピークメモリや実行時間の測定結果を保証するテストではありません。

`ToolsList_SuggestionDisclosesConditionalGitHubPublication_Issue5396` は共通の
フィクスチャで、既定・compact・full・ツール選択・ページ付きの一覧を検証します。
送信ツールを呼び出さず、外部サービスとのやり取りを示す注釈、その他の注釈値、
Expand Down
19 changes: 19 additions & 0 deletions changelog.d/unreleased/5415.fixed.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
---
category: fixed
issues:
- 5415
affected:
- src/CodeIndex/Database/IndexedFindQuerySource.cs
- src/CodeIndex/Database/IndexedFindPipeline.cs
- tests/CodeIndex.Tests/FindChunkReadTests.cs
- docs/find-scan-controls.md
- TESTING_GUIDE.md
---

## English

- **Line-local find no longer sorts source content (#5415)** — literal and line-regex queries read chunks in index order, with bounded metadata pages for older indexes. Source content is fetched separately per chunk, preserving overlap ownership, missing-source behavior, counts, scan limits and continuation without requiring reindexing.

## 日本語

- **行単位の find で本文をソートしなくなりました (#5415)** — リテラル・行単位正規表現検索は索引順にチャンクを読み、旧索引では上限付きのメタデータページを使います。本文はチャンクごとに別途取得し、重複行の優先順、本文欠落時の挙動、件数、走査上限、継続取得を維持します。再索引は不要です。
21 changes: 21 additions & 0 deletions docs/find-scan-controls.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,17 @@

## English

Literal and line-regex `find` read indexed chunks in line/chunk order without
sorting source content. Current indexes use the ordered partial chunk index;
older indexes page at most 64 scalar chunk records through a metadata-only sort,
then read each selected chunk's content separately. Pages use the last ordering
key rather than an increasing offset. Retained sort data is bounded, but legacy
pages may rescan file metadata and one chunk's content is still read at a time.
A bounded-memory NULL-content check preserves the existing missing-content error
by selecting the all-row fallback for such files; that check may scan the file's
chunk entries. Missing chunks never trigger live-file reads. Overlap ownership,
counts, line caps and continuation positions are unchanged; no reindex is required.

### CLI cursor validation errors (#5412)

Bounded `find` validation honors machine output selected by `--json`,
Expand Down Expand Up @@ -222,6 +233,16 @@ text or JSON output when context from `--before`, `--after`, or

## 日本語

リテラル検索と行単位の正規表現 `find` は、本文をソートせず行・チャンク順に索引を読みます。
現在の索引では順序付き部分インデックスを使い、旧索引では最大64件のチャンクメタデータだけを
ソートしてから、選択した各チャンクの本文を個別に取得します。ページは増大する offset ではなく
直前の順序キーから再開します。ソートの保持量には上限がありますが、旧索引ではページごとに
ファイルのメタデータを再走査する場合があり、本文もチャンク1個単位では読み取ります。
NULL 本文の確認は保持メモリを制限して行い、該当するファイルでは全行を対象とする代替経路により
従来の本文欠落エラーを維持します。この確認もファイルのチャンク行を走査する場合があります。
チャンクが欠けても実ファイルへ読み取りを切り替えません。重複行の優先順、件数、行上限、
継続位置は従来どおりで、再索引は不要です。

### CLI カーソルの検証エラー (#5412)

上限付き `find` の検証は、`--json`、`--json=ndjson`、`--format json`、
Expand Down
2 changes: 1 addition & 1 deletion src/CodeIndex/Database/IndexedFindPipeline.cs
Original file line number Diff line number Diff line change
Expand Up @@ -139,7 +139,7 @@ private bool ScanFileLines(
var origins = request.SemanticFilters is null ? null : _owner.CreateFindOriginContext(file, request.CancellationToken);
var firstContextLine = Math.Max(1, file.FirstEligibleLine - collector.ContextBefore);
var stopScanning = false;
foreach (var indexedLine in _querySource.EnumerateIndexedFileLines(file.Id))
foreach (var indexedLine in _querySource.EnumerateIndexedFileLines(file.Id, request.CancellationToken))
{
request.CancellationToken.ThrowIfCancellationRequested();
if (indexedLine.Number > file.TotalLines)
Expand Down
92 changes: 71 additions & 21 deletions src/CodeIndex/Database/IndexedFindQuerySource.cs
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,18 @@ namespace CodeIndex.Database;

public partial class DbReader
{
internal const int FindChunkPageSize = 64;

internal static string FindLineChunkPageSql(bool ordered, bool resumed) => $"""
SELECT c.id, c.start_line, c.end_line, c.chunk_index
FROM chunks c{(ordered ? $" INDEXED BY {BoundedResourceReadChunkIndexName}" : "")}
WHERE c.file_id = @fileId
{(ordered ? "AND c.content IS NOT NULL" : "")}
{(resumed ? "AND (c.start_line, c.chunk_index) > (@startLine, @chunkIndex)" : "")}
ORDER BY c.start_line, c.chunk_index
LIMIT {FindChunkPageSize}
""";

private sealed class IndexedFindQuerySource(DbReader owner)
{
private readonly DbReader _owner = owner;
Expand Down Expand Up @@ -85,37 +97,75 @@ SELECT DISTINCT find_chunk.file_id
return command;
}

internal IEnumerable<IndexedLine> EnumerateIndexedFileLines(long fileId)
internal IEnumerable<IndexedLine> EnumerateIndexedFileLines(long fileId, CancellationToken cancellationToken)
{
// The partial index excludes NULL source. Keep the original read failure for
// such rows by routing that file through the all-row metadata fallback.
// 部分索引から除外される NULL 本文は、全行のメタデータ経路で従来の読取エラーを維持する。
var ordered = _owner._chunkIndexes.Contains(BoundedResourceReadChunkIndexName)
&& !HasNullChunkContent(fileId);
using var pageCommand = _owner._conn.CreateCommand();
SqliteCommandPolicy.Add(pageCommand, "@fileId", fileId);
var startLine = pageCommand.Parameters.AddWithValue("@startLine", 0);
var chunkIndex = pageCommand.Parameters.AddWithValue("@chunkIndex", 0);
using var command = _owner._conn.CreateCommand();
command.CommandText = @"
SELECT c.start_line, c.end_line, c.content
FROM chunks c
WHERE c.file_id = @fileId
ORDER BY c.start_line, c.chunk_index";
SqliteCommandPolicy.Add(command, "@fileId", fileId);
command.CommandText = "SELECT content FROM chunks WHERE id = @chunkId";
var chunkId = command.Parameters.AddWithValue("@chunkId", 0L);

var lastEmittedLine = 0;
using var reader = command.ExecuteTrackedReader();
while (reader.TrackedRead())
var resumed = false;
while (true)
{
var chunkStartLine = reader.GetInt32(0);
var chunkEndLine = reader.GetInt32(1);
var chunkLines = reader.GetString(2).Split('\n');
var lineCount = chunkEndLine - chunkStartLine + 1;

for (var index = 0; index < chunkLines.Length && index < lineCount; index++)
cancellationToken.ThrowIfCancellationRequested();
// LIMIT bounds the legacy top-N sort to scalar metadata. Source is loaded
// only after selection, so neither SQLite nor the page retains file content.
// 旧索引のソートを少量のメタデータに制限し、選択後にだけ本文を取得する。
pageCommand.CommandText = FindLineChunkPageSql(ordered, resumed);
var count = 0;
using (var page = pageCommand.ExecuteTrackedReader())
{
var absoluteLine = chunkStartLine + index;
if (absoluteLine <= lastEmittedLine)
continue;

lastEmittedLine = absoluteLine;
yield return new IndexedLine(absoluteLine, chunkLines[index]);
while (page.TrackedRead())
{
cancellationToken.ThrowIfCancellationRequested();
count++;
var chunkStartLine = page.GetInt32(1);
var chunkEndLine = page.GetInt32(2);
startLine.Value = chunkStartLine;
chunkIndex.Value = page.GetInt32(3);
chunkId.Value = page.GetInt64(0);
using var reader = command.ExecuteTrackedReader();
if (!reader.TrackedRead())
throw new InvalidOperationException("Indexed find chunk is no longer available.");
var chunkLines = reader.GetString(0).Split('\n');
var lineCount = chunkEndLine - chunkStartLine + 1;

for (var index = 0; index < chunkLines.Length && index < lineCount; index++)
{
cancellationToken.ThrowIfCancellationRequested();
var absoluteLine = chunkStartLine + index;
if (absoluteLine <= lastEmittedLine)
continue;

lastEmittedLine = absoluteLine;
yield return new IndexedLine(absoluteLine, chunkLines[index]);
}
}
}
if (count < FindChunkPageSize)
yield break;
resumed = true;
}
}

private bool HasNullChunkContent(long fileId)
{
using var command = _owner._conn.CreateCommand();
command.CommandText = "SELECT 1 FROM chunks WHERE file_id = @fileId AND content IS NULL LIMIT 1";
SqliteCommandPolicy.Add(command, "@fileId", fileId);
using var reader = command.ExecuteTrackedReader();
return reader.TrackedRead();
}

internal static IEnumerable<FindLineMatch> EnumerateLineMatches(
string lineText,
string? fileLang,
Expand Down
Loading
Loading