XQuery
Full-Text Search
XQuery Full-Text β the contains text clause, match options, and phx: functions
#Full-Text Search
XQuery's built-in contains() function does exact substring matching. It finds "data" inside "database" but cannot search linguistically. XQuery Full Text adds language-aware matching: stemming, case-insensitive matching, and relevance scoring, evaluated as part of the query itself.
Important β verify before relying on this page. This page was corrected against the engine on
mainas of 2026-09-14. Two things are worth knowing before you use any of it:
The entry point is the
contains textclause, not a function. Earlier revisions of this page described anft:contains($node, "term")function call. No such function has ever existed β compiling a query that calls it fails withUnknown function: contains#2. The real syntax is the W3C XQuery Full Textcontains textclause shown below.
contains textfails at query-compile time beforePhoenixmlDb.XQuery1.8.0. On earlier releases everycontains textquery β with or without match options β throws aNullReferenceExceptionfromPhoenixmlDb.XQuery.Analysis.SchemaFeatureChecker.VisitStepExpression, before the query runs. A plain, non-full-text predicate compiles and runs fine, so the failure is specific tocontains text. Tracked asphoenixmldb-xquery#15and fixed in 1.8.0, which the engine now pins. On 1.8.0 and later the clause compiles and evaluates; verified by running the examples on this page through the publishedxquery41.8.0 tool.
#contains text β The Basic Clause
contains text is an XQuery expression, not a function: an operand (the node or nodes to search) on the left, the clause, then a full-text selection (what to look for) on the right.
(: Search the description element for "database" :)
//book[description contains text "database"]
(: Search ALL text content of the book element :)
//book[. contains text "xml query"]
(: Search with match options :)
//book[title contains text "xml" using stemming using case insensitive]
#Searching Multiple Fields
(: Search title OR description :)
//book[title contains text "xml" or description contains text "xml"]
#Using contains text in FLWOR Expressions
for $article in //article
where $article/body contains text "machine learning"
order by phx:score($article/body) descending
return
<result>
<title>{ $article/title/text() }</title>
<score>{ phx:score($article/body) }</score>
</result>
phx:score takes the node, not the search term β see phx:score(), below, for why.
This example does not produce useful output on 2.0.0.
phx:scorereturns0.0for every node on this release, so the ordering is arbitrary and every<score>is0. Re-measured on the publishedxquery42.0.0 tool: acontains textmatch that succeeds still scores0. The shape is correct; the scores are not yet.
#Match Options
Match options follow the search string and control how matching is performed, combined with successive using clauses. The grammar accepts more options than the engine currently acts on β each subsection below says which.
#Language
//article[. contains text "running" using language "en"]
Affects stemming rules and tokenization. Functional.
#Stemming
(: Without stemming β only matches literal "running" :)
//article[. contains text "running"]
(: With stemming β matches "run", "runs", "running", "ran" :)
//article[. contains text "running" using stemming]
Functional.
#Case Sensitivity
(: Default: case insensitive β matches "XML", "xml", "Xml" :)
//doc[title contains text "xml"]
(: Case sensitive β only matches exact case :)
//doc[title contains text "XML" using case sensitive]
Functional.
#Diacritics, Wildcards, Stop Words, and Thesaurus β parsed, not applied
The grammar also accepts using diacritics sensitive/insensitive, using wildcards/no wildcards, using stop words (...)/using no stop words, and using thesaurus "file". All four parse without error and are carried into the query's AST β but none of them currently reach the analyzer that does the actual matching. A query like:
//doc[name contains text "cafe" using diacritics sensitive]
//doc[. contains text "data" using wildcards]
//doc[. contains text "the art of war" using stop words ("the", "of")]
//doc[. contains text "fast" using thesaurus "thesaurus.xml"]
compiles but behaves exactly as if the using clause were absent: diacritics are always folded,
no glob expansion happens, no synonym is added, and the analyzer's own stop-word handling is
unaffected by what you wrote. Treat these four as accepted-but-inert until the underlying
analyzer is wired up.
using no stop wordsdoes not give you exact phrase matching. The analyzer removes stop words regardless, and the option does not stop it. Re-measured on 2.0.0,. contains text 'walrus carpenter'matches<p>the walrus and the carpenter</p>with and withoutusing no stop wordsβ identical results. If you reach for this option to make a phrase position-exact against the source text, it will silently not do that. See why the two phrase matchers differ.
Two syntax forms that look plausible are not supported at all β they fail to parse:
-
There is no
using stop words default. The only stop-word forms the grammar accepts are an explicit list (using stop words ("the", "a", "an")) orusing no stop words. There is no keyword for "use the language's built-in list" β and per above, even the explicit-list and no-stop-words forms don't currently change matching. -
There is no
using thesaurus at "file" relationship "type". The grammar accepts exactly one string literal:using thesaurus "thesaurus.xml".atandrelationshipare not part of it.
#Combining Match Options
Options are composable at the grammar level:
//article[body contains text "running"
using stemming
using case insensitive
using language "en"]
#Positional Filters
Positional filters constrain where and how search terms appear relative to each other, and follow the full-text selection (not the match options):
(: "introduction" must appear before "conclusion" :)
//doc[. contains text ("introduction" ftand "conclusion") ordered]
(: "xml" and "database" within 5 words of each other :)
//doc[. contains text ("xml" ftand "database") window 5 words]
The grammar also defines distance N words, same sentence, same paragraph, at start, at end, and entire content. As with the match options above, this page has not verified which of these actually change matching versus parse-and-ignore β the contains text compile failure blocks testing all of them the same way. Confirm behavior against your own build before depending on a specific filter.
#Logical Combinations
ftand, ftor, and ftnot combine search conditions inside a single full-text selection β they are not XPath's and/or, and don't need a repeated contains text:
(: Document must contain both "xml" and "database" :)
//doc[. contains text ("xml" ftand "database")]
(: Document contains "xml" or "json" :)
//doc[. contains text ("xml" ftor "json")]
(: Contains "database" but NOT "relational" :)
//doc[. contains text ("database" ftand ftnot "relational")]
ftnot is a unary prefix, not a binary infix. Writing ("database" ftnot "relational")
is a parse error β XPST0003: mismatched input 'ftnot' expecting ')'. Combine it with ftand
as above.
#Full-Text Functions
These are ordinary functions in https://schemas.phoenixml.dev/2026/functions β unlike
contains text, they use normal function-call syntax.
phxis predeclared. You do not declare it. FromPhoenixmlDb.XQuery2.0.0 the engine bindsphxon every query path, so the calls below work with no prolog. A container'sDefaultNamespacesand a query's owndeclare namespacestill take precedence, but a host can no longer rebindphxto a different URI β that is a compile error.
Renamed in 2.0.0 β the old names no longer compile. Releases before 2.0.0 put these functions in
http://www.w3.org/2007/xpath-full-textwith the prefixft, which had to be declared. That namespace is retired and there are no aliases, so a query using it fails at compile time rather than behaving differently.
before 2.0.0
2.0.0 and later
phx:stem
phx:stem
phx:tokenize
phx:tokenize
phx:score
phx:score
phx:is-stop-word
phx:is-stop-word
phx:thesaurus-lookup
phx:thesaurus-lookup
dbxml:metadata
phx:metadataThe old namespace was the W3C's own β used by the Full Text specification for its schemas β so it carried a false suggestion that these were standard, portable functions. They never were, and they are not now: no other XQuery processor provides them.
Provenance. The examples in this section were not covered by this page's original sample
verification. They were re-verified against engine 4231a6b with PhoenixmlDb.XQuery 1.8.0 on
2026-09-14, and re-run under the phx names against the published xquery4 2.0.0 on
2026-09-15. The signatures and outputs below are what those runs produced.
#phx:stem()
phx:stem($term as xs:string) as xs:string
phx:stem($term as xs:string, $language as xs:string) as xs:string
phx:stem("running", "en") (: "run" :)
#phx:tokenize()
phx:tokenize($text as xs:string?) as xs:string*
phx:tokenize($text as xs:string?, $language as xs:string) as xs:string*
Breaks text into tokens using the same analyzer contains text uses β which is what makes it
useful: it shows you the stream your phrase queries are actually matched against.
phx:tokenize("Hello, world! This is a test.")
(: ("hello", "world", "test") :)
Tokens are lower-cased, and stop words are dropped. This, is and a do not survive. If a
phrase query is matching more than you expect, running the text through phx:tokenize will usually
show you why β see
why the two phrase matchers differ.
#phx:is-stop-word()
phx:is-stop-word($word as xs:string) as xs:boolean
One argument, not two. A two-argument call fails with XPST0017: Unknown function: is-stop-word#2.
Re-measured on the published xquery4 2.0.0 tool, the result is sharper than "inaccurate":
the function returns false for every stop word and true only for input containing no letters.
|
call |
result |
call |
result |
|
|---|---|---|---|---|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
Meanwhile phx:tokenize("the quick brown fox jumps over the lazy dog") returns
quick|brown|fox|jumps|over|lazy|dog β both instances of the removed. This function and the
analyzer do not agree, so do not use it to predict what phx:tokenize or contains text will
do. In effect it tests "produced no letters", not "is a stop word". Tracked as
phoenixmldb-xquery#70.
On 2.0.0 this is a wiring bug, not a difference of configuration. The two paths select
different built-in analyzers by accident: phx:is-stop-word analyzes with stemming off, which
picks an analyzer that has no stop-word filter at all, so every word produces a token and the
answer is always false. contains text and phx:tokenize run with stemming on, which picks the
English analyzer, and that one does remove stop words.
#phx:score()
phx:score($node as node()) as xs:double
Takes one argument β the node β not the node and a search term.
Scores are
0.0on 2.0.0.phx:scorereturns0.0even immediately after a matchingcontains texton the same node, so awhere $score > 0filter returns nothing at all. Treat scoring as not working on this release. Tracked asphoenixmldb-xquery#71; not a documentation defect.
#phx:thesaurus-lookup()
phx:thesaurus-lookup($term as xs:string) as xs:string*
phx:thesaurus-lookup($term as xs:string, $relationship as xs:string) as xs:string*
The term comes first. An earlier revision of this page documented
phx:thesaurus-lookup($thesaurus, $term), taking a thesaurus file as the first argument. Called
that way it returns an empty sequence.
#Practical Examples
Every example in this section that uses
phx:scoreis shape-correct and does not work on 2.0.0. Scores come back0.0for every node, so anorder byon them is arbitrary and awhere $score > 0filter returns nothing at all. They are kept because the query shape is right and will start working when scoring does. No prolog declaration is needed:phxis predeclared.
#Document Search with Scoring
declare variable $query external;
for $doc in collection("documents")
where $doc contains text { $query } using stemming using language "en"
let $score := phx:score($doc)
where $score > 0
order by $score descending
return
<result score="{ $score }">
<title>{ $doc//title/text() }</title>
</result>
Note the { $query } form: a full-text selection can be a computed string ({ expr }), not only a string literal, so the search term can come from a variable.
#Content Management β Search and Highlight
declare function local:search-articles(
$terms as xs:string,
$max-results as xs:integer
) as element(results) {
let $matches :=
for $article in collection("cms")/article
where $article/body contains text { $terms } using stemming using case insensitive using language "en"
let $score := phx:score($article/body)
order by $score descending
return $article
return
<results total="{ count($matches) }">
{
for $m in subsequence($matches, 1, $max-results)
return
<article id="{ $m/@id }">
<title>{ $m/title/text() }</title>
</article>
}
</results>
};
local:search-articles("machine learning neural networks", 10)
#Multi-Language Search
declare function local:search(
$collection as xs:string,
$terms as xs:string,
$lang as xs:string
) as element()* {
for $doc in collection($collection)
where $doc contains text { $terms } using stemming using language { $lang }
let $score := phx:score($doc)
order by $score descending
return $doc
};
(: English search β "running" matches "run" :)
local:search("articles-en", "running databases", "en")
#Full-text index acceleration
On 1.8.0 and later, contains text evaluates by scanning; it does not consult PhoenixmlDb's Lucene-backed full-text index. That index exists (see Full-Text Search in the PhoenixmlDb section) but is reached through IndexManager.SearchFullText, a separate C# entry point β not through this XQuery clause. Accelerating contains text with that index is planned but not built, and has a documented prerequisite: the index matches phrases more strictly than contains text itself does, so using it as a naive candidate source would silently drop matches the scanning evaluator would otherwise confirm.
#See also
-
PhoenixmlDb Full-Text Search β the Lucene-backed index this page's
contains textdoes not (yet) use