Skip to content
Open
22 changes: 15 additions & 7 deletions docs/reference/core.md
Original file line number Diff line number Diff line change
Expand Up @@ -65,20 +65,28 @@ specify init my-project --integration copilot --preset compliance
## Naming Features with the Helper Scripts

When calling the bundled `create-new-feature` helper scripts directly, generated
names retain only ASCII letters and digits. A description entirely in a non-Latin
script, or made only of punctuation, can therefore produce an empty suffix such
as `001-`. The scripts warn on stderr when this happens, including during a dry
run; JSON output remains parseable.
names retain Unicode letters and decimal digits in UTF-8, so a description such as
`添加用户` produces `001-添加用户`. Descriptions made only of punctuation can still
produce an empty suffix such as `001-`; the scripts warn on stderr when this
happens, including during a dry run. JSON output remains parseable.

Keep the original description and supply a readable ASCII short name:
To choose a different name, keep the original description and supply a short name:

```bash
bash .specify/scripts/bash/create-new-feature.sh --json --short-name user-auth "添加用户"
bash .specify/scripts/bash/create-new-feature.sh --json --short-name 用户管理 "添加用户"
```

The Python helper also accepts `--short-name`; the PowerShell helper uses
`-ShortName`. A supplied short name is cleaned by the same rules, so it must
contain at least one ASCII letter or digit.
contain at least one letter or digit. For non-ASCII names, the Bash helper needs
an installed UTF-8 locale and a Python 3 interpreter for Unicode classification.
ASCII input, including tabs and newlines, is sanitized without either requirement.
If `LC_ALL` is non-empty, Bash uses that locale rather than selecting another:
Unicode names fail with an error if the selected locale is not usable for UTF-8
names. With `LC_ALL` unset or empty, Bash selects an installed UTF-8 locale even
when `LANG` or `LC_CTYPE` names a non-UTF-8 locale.
ASCII capitals are lowercased; non-ASCII letter casing is preserved across the
script variants.

## Check Installed Tools

Expand Down
2 changes: 1 addition & 1 deletion scripts/bash/common.sh
Original file line number Diff line number Diff line change
Expand Up @@ -112,7 +112,7 @@ read_feature_json_feature_directory() {
fi
if [[ -z "$_fd" ]] && command -v python3 >/dev/null 2>&1; then
# Use Python so pretty-printed/multi-line JSON still parses correctly.
if ! _fd=$(python3 -c "import json,sys; d=json.load(open(sys.argv[1])); v=d.get('feature_directory'); print(v if v else '')" "$fj" 2>/dev/null); then
if ! _fd=$(python3 -c "import json,sys; d=json.load(open(sys.argv[1], encoding='utf-8')); v=d.get('feature_directory'); print(v if v else '')" "$fj" 2>/dev/null); then
_fd=''
fi
fi
Expand Down
119 changes: 98 additions & 21 deletions scripts/bash/create-new-feature.sh
Original file line number Diff line number Diff line change
Expand Up @@ -146,17 +146,82 @@ spec_prefix_exists() {

# Function to clean and format a branch name
#
# Three details keep this byte-identical to the Python and PowerShell twins:
# * LC_ALL=C -- in a UTF-8 locale glibc resolves the a-z *range* through
# collation, so [^a-z0-9] keeps accented lowercase letters that
# re.sub(r"[^a-z0-9]", ...) and .NET's -replace both strip.
# Three details keep this consistent with the Python and PowerShell twins:
# * Unicode classification uses Python: POSIX [:alnum:] differs by platform.
# * `--*` instead of the GNU-only `\+`, which POSIX/BSD sed reads as a literal
# '+', leaving repeated separators uncollapsed on macOS.
# * printf instead of echo, so a name of "-n"/"-e"/"-E" is text, not options.
contains_non_ascii() {
LC_ALL=C grep -q '[^[:print:][:cntrl:]]'
}

UNICODE_LOCALE=""
locale_candidates=(C.UTF-8 C.utf8 en_US.UTF-8 en_US.utf8 "${LC_CTYPE:-${LANG:-}}")
if [ -n "${LC_ALL:-}" ]; then
locale_candidates=("$LC_ALL")
fi
for candidate in "${locale_candidates[@]}"; do
if [ -n "$candidate" ] && [ "$(printf 'é。' | LC_ALL="$candidate" sed 's/[^[:alnum:]]/-/g' 2>/dev/null)" = 'é-' ]; then
UNICODE_LOCALE="$candidate"
break
fi
done

if [ -z "$UNICODE_LOCALE" ]; then
UNICODE_LOCALE=C
if printf '%s' "${SHORT_NAME:-$FEATURE_DESCRIPTION}" | contains_non_ascii; then
if [ -n "${LC_ALL:-}" ]; then
echo "Error: A UTF-8 locale is required to create a Unicode feature name; LC_ALL=$LC_ALL is not usable" >&2
else
echo "Error: A UTF-8 locale is required to create a Unicode feature name" >&2
fi
exit 1
Comment thread
Copilot marked this conversation as resolved.
fi
fi

unicode_words() {
local name="${1//$'\n'/ }"
local separator="$2"
if printf '%s' "$name" | contains_non_ascii; then
local -a python_cmd=()
local override="${SPECKIT_PYTHON_EXECUTABLE:-${SPECKIT_PYTHON:-}}"
if [ -n "$override" ] && command -v "$override" >/dev/null 2>&1 &&
"$override" -c 'import sys; raise SystemExit(sys.version_info.major != 3)' >/dev/null 2>&1; then
python_cmd=("$override")
else
local python_line
while IFS= read -r python_line; do
python_cmd+=("$python_line")
done < <(_python3_command)
fi
if [ "${#python_cmd[@]}" -eq 0 ]; then
echo "Error: Python 3 is required to create a Unicode feature name" >&2
return 1
fi
printf '%s' "$name" | "${python_cmd[@]}" -c '
import sys
value = sys.stdin.buffer.read().decode("utf-8")
lower = str.maketrans("ABCDEFGHIJKLMNOPQRSTUVWXYZ", "abcdefghijklmnopqrstuvwxyz")
result = "".join(
char if char.isalpha() or char.isdecimal() else sys.argv[1]
for char in value.translate(lower)
)
sys.stdout.buffer.write(result.encode("utf-8"))
' "$separator"
else
printf '%s' "$name" | LC_ALL=C tr '[:upper:]' '[:lower:]' | LC_ALL=C sed "s/[^a-z0-9]/$separator/g"
fi
}

clean_branch_name() {
local name="$1"
local -x LC_ALL=C
printf '%s\n' "$name" | tr '[:upper:]' '[:lower:]' | sed 's/[^a-z0-9]/-/g' | sed 's/--*/-/g' | sed 's/^-//' | sed 's/-$//'
local cleaned
cleaned=$(unicode_words "$name" '-') || return 1
printf '%s\n' "$cleaned" | sed 's/--*/-/g' | sed 's/^-//' | sed 's/-$//'
}

branch_byte_count() {
printf '%s' "$1" | wc -c | tr -d '[:space:]'
}

# Fit a feature prefix and suffix within GitHub's branch-name limit.
Expand All @@ -165,11 +230,25 @@ fit_branch_name() {
local branch_suffix="$2"
local branch_name="${feature_num}-${branch_suffix}"

if [ ${#branch_name} -gt $MAX_BRANCH_LENGTH ]; then
if [ "$(branch_byte_count "$branch_name")" -gt "$MAX_BRANCH_LENGTH" ]; then
local prefix_length=$(( ${#feature_num} + 1 ))
local max_suffix_length=$((MAX_BRANCH_LENGTH - prefix_length))
local truncated_suffix
truncated_suffix=$(printf '%s' "$branch_suffix" | cut -c "1-$max_suffix_length" | sed 's/-$//')
local -x LC_ALL="$UNICODE_LOCALE"
local low=0 high=${#branch_suffix} mid
if (( high > max_suffix_length )); then
high=$max_suffix_length
fi
while (( low < high )); do
mid=$(((low + high + 1) / 2))
if [ "$(branch_byte_count "${branch_suffix:0:$mid}")" -le "$max_suffix_length" ]; then
low=$mid
else
high=$((mid - 1))
fi
done
truncated_suffix="${branch_suffix:0:$low}"
truncated_suffix="${truncated_suffix%-}"
branch_name="${feature_num}-${truncated_suffix}"
fi

Expand Down Expand Up @@ -209,27 +288,25 @@ generate_branch_name() {
# Common stop words to filter out
local stop_words="^(i|a|an|the|to|for|of|in|on|at|by|with|from|is|are|was|were|be|been|being|have|has|had|do|does|did|will|would|should|could|can|may|might|must|shall|this|that|these|those|my|your|our|their|want|need|add|get|set)$"

# Convert to lowercase and split into words. LC_ALL=C for the same
# collation reason documented on clean_branch_name, and so the `grep -qw`
# acronym probe below uses ASCII word boundaries like the Python twin's
# (?<![0-9A-Za-z_]) lookarounds.
local -x LC_ALL=C
local clean_name=$(printf '%s' "$description" | tr '[:upper:]' '[:lower:]' | sed 's/[^a-z0-9]/ /g')
# Use a UTF-8 locale for character-safe length checks and split words.
local -x LC_ALL="$UNICODE_LOCALE"
local clean_name
clean_name=$(unicode_words "$description" ' ') || return 1

# Filter words: remove stop words and words shorter than 3 chars (unless they're uppercase acronyms in original)
local meaningful_words=()
for word in $clean_name; do
# Skip empty words
[ -z "$word" ] && continue

# Keep words that are NOT stop words AND (length >= 3 OR are potential acronyms)
if ! echo "$word" | grep -qiE "$stop_words"; then
if [ ${#word} -ge 3 ]; then
# Retain non-ASCII words even when shorter than three characters.
if ! printf '%s\n' "$word" | LC_ALL=C grep -qE "$stop_words"; then
if [ ${#word} -ge 3 ] || printf '%s' "$word" | contains_non_ascii; then
meaningful_words+=("$word")
# Keep short words that appear as an uppercase acronym in the original.
# Uppercase via tr and match with grep -w (both portable) rather than
# bash's 4+ "^^" case expansion (breaks on macOS bash 3.2) and \b (non-POSIX).
elif printf '%s' "$description" | grep -qw -- "$(printf '%s' "$word" | tr '[:lower:]' '[:upper:]')"; then
elif printf '%s' "$description" | LC_ALL=C grep -qw -- "$(printf '%s' "$word" | LC_ALL=C tr '[:lower:]' '[:upper:]')"; then
meaningful_words+=("$word")
fi
fi
Expand Down Expand Up @@ -266,7 +343,7 @@ else
fi

if [ -z "$BRANCH_SUFFIX" ]; then
echo "[specify] Warning: Feature name is empty after removing unsupported characters. Use --short-name with ASCII letters or digits (for example, user-auth)." >&2
echo "[specify] Warning: Feature name is empty after removing unsupported characters. Use --short-name with letters or digits (for example, user-auth)." >&2
fi

# Warn if --number and --timestamp are both specified
Expand Down Expand Up @@ -339,8 +416,8 @@ ORIGINAL_BRANCH_NAME="${FEATURE_NUM}-${BRANCH_SUFFIX}"
BRANCH_NAME=$(fit_branch_name "$FEATURE_NUM" "$BRANCH_SUFFIX")
if [ "$BRANCH_NAME" != "$ORIGINAL_BRANCH_NAME" ]; then
>&2 echo "[specify] Warning: Branch name exceeded GitHub's 244-byte limit"
>&2 echo "[specify] Original: $ORIGINAL_BRANCH_NAME (${#ORIGINAL_BRANCH_NAME} bytes)"
>&2 echo "[specify] Truncated to: $BRANCH_NAME (${#BRANCH_NAME} bytes)"
>&2 echo "[specify] Original: $ORIGINAL_BRANCH_NAME ($(branch_byte_count "$ORIGINAL_BRANCH_NAME") bytes)"
>&2 echo "[specify] Truncated to: $BRANCH_NAME ($(branch_byte_count "$BRANCH_NAME") bytes)"
fi

FEATURE_DIR="$SPECS_DIR/$BRANCH_NAME"
Expand Down
58 changes: 40 additions & 18 deletions scripts/powershell/create-new-feature.ps1
Original file line number Diff line number Diff line change
Expand Up @@ -83,10 +83,32 @@ function Test-SpecPrefixInUse {
Select-Object -First 1)
}

function ConvertTo-AsciiLower {
param([string]$Name)

return [regex]::Replace($Name, '[A-Z]', { param($match) $match.Value.ToLowerInvariant() })
}

function ConvertTo-UnicodeWords {
param([string]$Name, [string]$Separator)

$lowerName = ConvertTo-AsciiLower -Name $Name
return [regex]::Replace($lowerName, '[\uD800-\uDBFF][\uDC00-\uDFFF]|[^\p{L}\p{Nd}]', {
param($match)
if ($match.Length -eq 2) {
$category = [System.Globalization.CharUnicodeInfo]::GetUnicodeCategory($match.Value, 0)
if ($category.ToString() -match '(Letter|DecimalDigitNumber)$') {
return $match.Value
}
}
return $Separator
})
}

function ConvertTo-CleanBranchName {
param([string]$Name)

return $Name.ToLower() -replace '[^a-z0-9]', '-' -replace '-{2,}', '-' -replace '^-', '' -replace '-$', ''
return (ConvertTo-UnicodeWords -Name $Name -Separator '-') -replace '-{2,}', '-' -replace '^-', '' -replace '-$', ''
}

function Get-FittedBranchName {
Expand All @@ -96,10 +118,15 @@ function Get-FittedBranchName {
)

$fittedName = "$FeatureNum-$BranchSuffix"
if ($fittedName.Length -gt $maxBranchLength) {
if ([System.Text.Encoding]::UTF8.GetByteCount($fittedName) -gt $maxBranchLength) {
$prefixLength = $FeatureNum.Length + 1
$maxSuffixLength = $maxBranchLength - $prefixLength
$truncatedSuffix = $BranchSuffix.Substring(0, [Math]::Min($BranchSuffix.Length, $maxSuffixLength))
$bytes = [System.Text.Encoding]::UTF8.GetBytes($BranchSuffix)
$bytesToUse = $maxSuffixLength
while ($bytesToUse -gt 0 -and ($bytes[$bytesToUse] -band 0xC0) -eq 0x80) {
$bytesToUse--
}
$truncatedSuffix = [System.Text.Encoding]::UTF8.GetString($bytes, 0, $bytesToUse)
$truncatedSuffix = $truncatedSuffix -replace '-$', ''
$fittedName = "$FeatureNum-$truncatedSuffix"
}
Expand Down Expand Up @@ -132,18 +159,19 @@ function Get-BranchName {
'want', 'need', 'add', 'get', 'set'
)

# Convert to lowercase and extract words (alphanumeric only)
$cleanName = $Description.ToLower() -replace '[^a-z0-9\s]', ' '
# Lowercase ASCII and extract Unicode words, matching the shell variant.
$cleanName = ConvertTo-UnicodeWords -Name $Description -Separator ' '
$words = $cleanName -split '\s+' | Where-Object { $_ }

# Filter words: remove stop words and words shorter than 3 chars (unless they're uppercase acronyms in original)
$meaningfulWords = @()
foreach ($word in $words) {
# Skip stop words
if ($stopWords -contains $word) { continue }
if ($stopWords -ccontains $word) { continue }

# Keep words that are length >= 3 OR appear as uppercase in original (likely acronyms)
if ($word.Length -ge 3) {
# Keep Unicode words even when short; ASCII words still need three
# characters or an uppercase acronym in the original.
if ($word.Length -ge 3 -or $word -match '[^\x00-\x7F]') {
$meaningfulWords += $word
} elseif ($Description -cmatch "(?<![0-9A-Za-z_])$($word.ToUpper())(?![0-9A-Za-z_])") {
# Keep short words only if they appear as uppercase in original (likely
Expand All @@ -164,13 +192,7 @@ function Get-BranchName {
} else {
# Fallback to original logic if no meaningful words found
$result = ConvertTo-CleanBranchName -Name $Description
# @() keeps this an array. ConvertTo-CleanBranchName blanks every
# non-[a-z0-9] character, so a description written in a non-Latin script
# (or made only of punctuation) leaves nothing for the pipeline to
# emit -- it yields $null, and [string]::Join on $null throws
# ArgumentNullException. With $ErrorActionPreference = 'Stop' that is
# terminating, so the script died with a .NET stack trace and exit 1
# where the bash and Python twins both return an empty suffix.
# @() keeps this an array when the description contains only separators.
$fallbackWords = @(($result -split '-') | Where-Object { $_ } | Select-Object -First 3)
return [string]::Join('-', $fallbackWords)
}
Expand All @@ -186,7 +208,7 @@ if ($ShortName) {
}

if (-not $branchSuffix) {
[Console]::Error.WriteLine("[specify] Warning: Feature name is empty after removing unsupported characters. Use -ShortName with ASCII letters or digits (for example, user-auth).")
[Console]::Error.WriteLine("[specify] Warning: Feature name is empty after removing unsupported characters. Use -ShortName with letters or digits (for example, user-auth).")
}

# Treat an explicit empty string as omitted, matching the bash and Python twins.
Expand Down Expand Up @@ -258,8 +280,8 @@ $originalBranchName = "$featureNum-$branchSuffix"
$branchName = Get-FittedBranchName -FeatureNum $featureNum -BranchSuffix $branchSuffix
if ($branchName -ne $originalBranchName) {
[Console]::Error.WriteLine("[specify] Warning: Branch name exceeded GitHub's 244-byte limit")
[Console]::Error.WriteLine("[specify] Original: $originalBranchName ($($originalBranchName.Length) bytes)")
[Console]::Error.WriteLine("[specify] Truncated to: $branchName ($($branchName.Length) bytes)")
[Console]::Error.WriteLine("[specify] Original: $originalBranchName ($([System.Text.Encoding]::UTF8.GetByteCount($originalBranchName)) bytes)")
[Console]::Error.WriteLine("[specify] Truncated to: $branchName ($([System.Text.Encoding]::UTF8.GetByteCount($branchName)) bytes)")
}

$featureDir = Join-Path $specsDir $branchName
Expand Down
Loading
Loading