After proceeding further in development, the only non-initFromReset initialization pattern that appears is constructing a raw instance, then immediately providing a computed tokenization. Thus, accepting the tokenization as an optional parameter will allow us a nice constructor spec for both init styles.
Note that this will temporarily cause predictive-text to throw out older context states (that may be valid rewind targets) in favor of newer ones (that are known to be invalid rewind targets) due to the current context-state caching logic. That'll be resolved later, in an upcoming PR - and preferably, this one shouldn't merge without that one.
This PR's changes aim to meet our current needs for the tracking of context tokenization (word boundaries) across edits... all while moving much closer to the design specified within the Correction-Search + Context-Tracking design doc.
Also places the class within its own separate source file.
Thanks to a PR review that caught an accidental restart that hid one of the Transforms
Co-authored-by: Eberhard Beilharz <ermshiperete@users.noreply.github.com>
Following from #14364, this PR integrates the new method with the main predictive-text context-tracking code, significantly reworking the `attemptMatchContext` method in the process. While further refactoring of the latter method is planned, this step allows us to verify that the new methods integrate properly with the main codebase in their current form.
This also comes with the benefit of simplifying `attemptMatchContext` _significantly_ - large parts of its code were refactored into `attemptTokenizedAlignment`, and the new logic patterns are generally more straightforward to parse and understand.
Following from #14363, this method performs context alignment calculations that may be
used to match forms of the context before and after an edit by aligning their tokens and
validating any edits that may have occurred.
Note that no 'tracked context' states are manipulated or altered by this method - it
solely calculates the alignment deltas needed to align the two contexts. Other methods
may then take these values and determine the edits that occurred during the associated
context transition as needed.
Note that the `attemptTokenizedAlignment` method is not integrated into the main codebase
for the predictive-text worker at this time. That said, this method _does_ integrate
the `isSubstitutionAlignable` method introduced by #14363.
This adds one new method within the predictive-text worker space: isSubstitutionAlignable. The method is designed to report whether or not two words are "related enough" to consider as an appropriate word-level "substitution" when matching the incoming context against previously-seen contexts - a process useful for facilitating delayed reversions, among other things.
It is not yet integrated with the main body of worker code, however.
This change marks the functions of `KeymanSentryManager` as public
or private, depending on whether or not they are used outside of the
module.
Test-bot: skip