Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 3 additions & 1 deletion Sources/PladderCore/Learning/CorrectionDiff.swift
Original file line number Diff line number Diff line change
Expand Up @@ -42,7 +42,9 @@ public enum CorrectionDiff {
/// two-word paste can still be corrected.
static let rewriteAllowance = 2
/// The alignment table's bound; past it the reading is skipped rather
/// than a quadratic table built. A 120 s dictation is far below it.
/// than a quadratic table built. A dictation of a few minutes is far
/// below it; one near the 10 min cap can pass it, and then nothing is
/// learned from that paste.
static let maximumCells = 4_000_000

/// Pairs in text order, empty when there is nothing to learn.
Expand Down
11 changes: 6 additions & 5 deletions Sources/PladderRefine/TranscriptPolisher.swift
Original file line number Diff line number Diff line change
Expand Up @@ -46,9 +46,10 @@ public actor TranscriptPolisher: TranscriptRefiner {
"""

/// Above this the transcript is split at sentence ends into windows of
/// about `chunkSize` words, each its own call. The context window is
/// about 4k tokens and a 120 s dictation is about 350 words, so this is
/// insurance rather than a path anyone takes.
/// about `chunkSize` words, each its own call with its own timeout. The
/// context window is about 4k tokens and a minute of speech is about 175
/// words, so a dictation past about three and a half minutes takes this
/// path; one near the 10 min cap is about six calls.
public static let chunkThreshold = 600
static let chunkSize = 300

Expand Down Expand Up @@ -173,8 +174,8 @@ public actor TranscriptPolisher: TranscriptRefiner {

// MARK: Chunking

/// The transcript as it is when it is short, which is every dictation
/// the 120 s cap allows; otherwise windows of about `chunkSize` words,
/// The transcript as it is when it is short, which is most dictations;
/// otherwise windows of about `chunkSize` words,
/// cut after a sentence end so no window starts mid-sentence.
static func chunks(of text: String) -> [String] {
let words = text.split(whereSeparator: \.isWhitespace)
Expand Down
Loading