Abstract
The theory of the ambiguity-resolved parameter significance (ARs) and ambiguity-resolved parameter normed (ARn) tests has been developed for global navigation satellite system mixed-integer model validation. This contribution describes how to apply these tests and evaluates their performance by analyzing their power as a function of bias magnitude and comparing them with the ambiguity-float (AF) and ambiguity-known significance tests. When testing for blunders in pseudoranges and carrier-phase cycle slips, the ARn and ARs tests perform similarly and generally outperform the AF test. When testing for ionosphere and troposphere delays, the ARs and ARn tests perform differently, and the ARs power function is spiky. We explain this behavior and show the performance of two methods in mitigating reductions in the ARs power function. This contribution demonstrates that the ARs and ARn tests can perform better the AF test, even when the ambiguity resolution success rate is not close to one.
1 INTRODUCTION
Modal validation is an essential step in global navigation satellite system (GNSS) data processing, which aims to ensure that the model to be used is sufficiently consistent with observations (Hofmann-Wellenhof et al., 2012; Teunissen & Montenbruck, 2017). The GNSS observation model can be misspecified by unmodeled effects, for instance, a blunder (also referred to as a measurement fault or an outlier) in a pseudorange observable, a carrier-phase cycle slip, and differential atmosphere delays that are not fully canceled out or corrected for (Braasch, 2017; Hernández-Pajares et al., 2011; Lawrence et al., 2006; Wanninger, 2004). The estimated parameters, e.g., user position coordinates, will be biased if misspecifications of the observation model remain unnoticed (Teunissen & Kleusberg, 2012). The detection, identification, and adaptation procedure (Baarda, 1968; Kok, 1984; Teunissen, 2018) is commonly used in model validation (Drevelle & Bonnifait, 2011; Hewitson & Wang, 2006; Perfetti, 2006; Yang et al., 2021). The overall model test can be used in the first step to detect whether the model is misspecified. In this step, only the default model under the null hypothesis is involved. If the detection test is not passed, an identification test can be applied in the second step to identify the discrepancy between the model and observed data. Multiple alternative hypotheses should be considered in this step. Finally, once the model misspecification has been identified, the model under the identified alternative hypothesis is used for data processing (Teunissen, 2000).
The GNSS observation model that involves precise carrier-phase observables contains integer-valued unknown carrier-phase cycle ambiguities (Leick et al., 2015). A theory on mixed-integer parameter inference has been established (Teunissen, 2017), while a theory on mixed-integer model validation remains under development. The mixed-integer validation problem is commonly transformed into a real-valued problem by (i) conducting model validation using the ambiguity-float (AF) solution (De Jong & Teunissen, 2000; El-Mowafy, 2014) or (ii) assuming the ambiguities to be known and ignoring the uncertainty of the resolved integer ambiguities (Feng et al., 2009; Wang et al., 2021; Zhang et al., 2023). Model validation cannot benefit from the integer constraints of the ambiguities in the first approach, and the second approach requires the ambiguity resolution success rate to be very close to one. Teunissen (2024) proposed an ambiguity-resolved (AR) detector, which utilizes the distribution of the resolved ambiguity and can outperform the AF detector in cases of a lower success rate (Yin et al., 2025). Teunissen (2025) developed a theory for the mixed-integer parameter significance test, in which two AR tests are introduced: the AR significance (ARs) test and the AR normed (ARn) test. The performance of the ARs and ARn tests for different model misspecifications has not yet been explored. This contribution describes how to apply the ARs and ARn tests and evaluates their performance by analyzing their power as a function of bias magnitude.
This contribution is organized as follows. Section 2 reviews the theory of ambiguity resolution and the probability mass function (PMF) of the integer ambiguity estimator, as well as the theory of parameter significance tests. Section 3 details the implementation of these methods and an evaluation of their power. A numerical experiment is conducted to provide insights into the required number of samples for a Monte Carlo simulation-based implementation of the ARs test. Section 4 presents the performance of the ARs and ARn tests in testing for a blunder in pseudorange, a carrier-phase cycle slip, and ionosphere and troposphere delays. Finally, Section 5 provides a summary and conclusions.
In this contribution, we use the following notations: E{·} and D{·} refer to the expectation and dispersion operator, respectively. The underlined symbol “” indicates a random variable. ℝm and ℤn denote the m-dimensional real space and the n-dimensional integer space. Nm(x, Q) represents an m-dimensional normal distribution with mean x and variance-covariance matrix Q. Ik denotes a k-dimensional identity matrix, 1k is a k × 1 vector with values of 1, and ui is a canonical vector with its i-th entry equal to one and all other entries equal to zero. refers to the best linear unbiased estimator inverse of matrix is the orthogonal projector that projects onto the range space of M, and projects onto its orthogonal complement. [⋅] refers to the rounding operation. ⊗ denotes the Kronecker product. represents the squared norm of a vector weighted by refers to the determinant of a matrix denotes the probability of an event. is the probability density function of a random variable .
2 AMBIGUITY-RESOLVED SIGNIFICANCE TEST
2.1 GNSS Mixed-Integer Observation Model
The parameter significance test aims to test for the significance of the entries of an unknown vector, i.e., to determine whether the tested entries should be included in the observation equations or can be considered small enough to be neglected. To perform the significance test, the linearized GNSS mixed-integer observation model can be written as follows (Leick et al., 2015; Teunissen & Montenbruck, 2017):
1
Here, A is the design matrix for the cycle ambiguities. B is the design matrix for the real-valued parameters that should be included in the default observation model, such as the (relative) position coordinates. C is the design matrix for the model parameters to be tested. We assume that the design matrix is of full column rank. contains the n-dimensional unknown integer-valued ambiguities, contains the real-valued unknowns, and c contains the model parameters to be tested. is the m-vector containing the pseudorange and carrier-phase observables, which are assumed to be normally distributed.
The parameter significance test serves to determine which of the following two hypotheses should be accepted:
2
and refer to the null and alternative hypotheses, respectively.
In this contribution, we utilize the short-baseline double-differenced (DD) observation model described below to evaluate the performance of the parameter significance tests. Similar performance can be expected for other models equivalent to this model, such as atmosphere-corrected network/precise point positioning real-time kinematic user models and short-baseline relative positioning models formulated with un/single-differenced observables (Odijk et al., 2016; Teunissen & Khodabandeh, 2015).
Let us assume that s satellites from a single constellation are tracked on f frequencies for k epochs. The design matrix of the ambiguity vector is written as follows:
3
where λj is the wavelength of frequency j. The design matrix of the baseline is written as follows:
4
Here, for a moving receiver, and for a static receiver. is an differencing matrix. G is an geometry matrix containing unit vectors from the satellites to the receiver, which is assumed to be constant as we focus on the single-epoch or short-time-span multi-epoch model in this contribution.
The elevation-weighted variance-covariance matrix for the DD observables is written as follows:
5
and are the zenith-referenced undifferenced pseudorange and carrier-phase standard deviations; we assume the pseudorange and carrier-phase observables to be uncorrelated. The matrix accounts for the time correlation between epochs (Odijk & Teunissen, 2008). With the identity matrix , we assume the observables from different frequencies to be uncorrelated and to have the same precision. is a diagonal matrix containing elevation-dependent weights (Eueler & Goad, 1991). For a satellite a elevation in degrees, the weight equals the following:
6
In this contribution, we evaluate the power of the parameter significance test for identifying (1) a blunder in pseudorange, (2) a carrier-phase cycle slip, (3) ionosphere delays, and (4) troposphere delays. The design matrices C corresponding to one-dimensional errors are given as follows. These matrices can be concatenated if multidimensional errors must be considered, e.g., for ionosphere delays that affect multiple satellites.
The design matrix for a blunder in the pseudorange observable of the v-th satellite at the j-th frequency and i-th epoch is written as follows:
7
where the size of a canonical vector ui is noted below the vector.
The observables from two epochs are used to detect a phase outlier. Let us assume that we identify a cycle slip in the phase observable of the second epoch from the v-th satellite on the j-th frequency. The corresponding design matrix is written as follows:
8
For an ionosphere delay that remains in the DD observables related to the v-th satellite on the i-th epoch, the corresponding design matrix is written as follows:
9
with:
10
For a troposphere zenith delay at the i-th epoch, the design matrix for the significance test is as follows:
11
Here, is the troposphere mapping function, and is used in our experiments.
2.2 Integer Ambiguity Estimator
Carrier-phase ambiguity resolution is the key step for conducting the AR tests for the GNSS mixed-integer observation model. This step starts with obtaining the float ambiguity estimator by neglecting the integer property of the ambiguities (Teunissen, 1995):
12
with . The float ambiguity vector can be resolved to an integer vector with one of the integer estimators (Teunissen, 1999b), such as the integer rounding (IR), integer bootstrapping (IB), or integer least-squares (ILS) estimator, denoted as , with . Before an integer estimator is applied to resolve the ambiguities, a decorrelation transformation should be applied to the float ambiguities, which can increase the success rate of the IR and IB estimators, as well as the efficiency of the ILS estimator (Teunissen, 1995; Verhagen et al., 2013). The ILS estimator provides the largest probability of resolving the ambiguities to the correct integer, i.e., the largest ambiguity resolution success rate (Teunissen, 1999a):
13
This corresponds to the integer vector nearest to the float ambiguity vector under the metric of , which can be efficiently obtained with the LAMBDA method (Massarweh et al., 2025; Teunissen, 1995).
The integer ambiguity estimator has discrete values, and different float ambiguity vectors can be resolved to the same integer vector; thus, I is a many-to-one mapping. The subset containing all of the float ambiguities that are resolved to the integer vector z is referred to as the pull-in region of (Teunissen, 2017).
The probability distribution of an integer ambiguity estimator is described by a PMF, which is required to obtain the probability distribution and critical value of the AR significance test. The PMF of an integer estimator results from the integration of the probability density function of the float ambiguity estimator over the pull-in regions of the integer estimator. Although the PMF of the integer estimator centers at the true but unknown ambiguity vector a, Equation (22) will show that only the shape of the PMF affects the distribution of the AR test statistics. This property allows us to obtain the PMF of an integer estimator with for conducting the AR significance test.
2.2.1 Integer Bootstrapping PMF
The PMF of the IB estimator can be obtained analytically through the following steps.
The PMF is defined over the infinite set , and we need to choose a finite set of integers for practical computation. The sum of the probability masses of the chosen integer set should be very close to one. The set can be defined as follows (Odolinski & Teunissen, 2022; Teunissen, 2005):
14
where is the quantile of the central chi-squared distribution with n degrees of freedom.
For each , the probability mass can be written as the product of n independent univariate probabilities (Teunissen, 2001):
15
where denotes the IB estimator. a is the true ambiguity vector, and can be used in computation. The conditional standard deviations , with I denoting , are the square roots of the diagonal values of the D matrix obtained from the triangular matrix factorization is the cumulative distribution function (CDF) of the standard normal distribution.
The success rate of the IB estimator follows from Equation (15):
16
2.2.2 Integer Least-Squares PMF
An analytical expression for the PMF of the ILS estimator is not available when the ambiguities are correlated. The steps for obtaining the ILS PMF through a Monte Carlo simulation are described below.
Generate -dimensional samples that follow , denoted as .
Compute the ILS solution for each sample, denoted as .
Count the number of samples that are fixed to the integer vector z, denoted as .
The PMF of the ILS estimator is approximated as follows:
17
and the empirical ILS success rate follows .
The standard deviation of a simulated probability mass is as follows:
18
The dashed lines in Figure 1 show the standard deviation as a function of the probability mass to be simulated, with (orange) and (yellow) samples. The left and right graphs show the simulation uncertainty for close to zero and unity, respectively. The probabilities of interest, and , are plotted in blue in the left and right graphs for comparison with the simulation uncertainty. With 106 samples, a probability mass of 10−4 can be simulated with a standard deviation of 10−5, which is 10% relative to 10−4. To simulate a 10−6 probability mass with the same relative standard deviation, 108 samples are required. The uncertainty here is related to the simulation of the probability mass. The impact of on simulating the critical value for the AR significance test will be discussed in the following section.
Standard deviation as a function of probability mass
Dashed lines show the standard deviation of the simulation as a function of the probability mass to be simulated for (orange) and (yellow) samples. The blue line shows the probability mass.
The success rate of the ILS estimator cannot be computed analytically (Teunissen, 1999a). In this contribution, we utilize the IB success rate (Equation (16)) as an approximation, which provides a tight and easy-to-compute lower bound for the ILS success rate (Massarweh et al., 2025; Verhagen, 2003).
2.3 Parameter Significance Tests
Parameter significance tests can be conducted based on the AF, ambiguity-known (AK), or AR estimators of the parameters c in Equation (1). We first introduce the three estimators and their distributions and then define the parameter significance tests based on these estimators.
When the ambiguity vector is unknown and we ignore its integer nature, the AF estimator of c is as follows:
19
If the ambiguity vector a is fully known, the AK estimator of c is then written as follows:
20
The ambiguity vector is commonly estimated rather than known. Although the AK estimator and the AK significance test cannot be strictly applied in practice, they are still useful for inclusion in the analysis, as they provide limiting cases of the AR estimator and AR significance test (Teunissen, 2025). By replacing the known ambiguity vector with the resolved ambiguity vector, we obtain the AR estimator of c:
21
The distributions of the three estimators under are written as follows (Teunissen, 2025):
22
The AF and AK estimators follow normal distributions, and the distribution of the AR estimator is a weighted sum of normal distributions, the means of which are shifted by . The weights are given by the PMF of the integer ambiguity estimator, and the shifts are computed with the integer vector z. The distributions under can be obtained with c = 0. Teunissen (2025) showed that and . The distribution of the AR real-valued parameter estimator was introduced by Teunissen (1998). Building on similar principles, Wu et al. (2008) and Khanafseh and Pervan (2010) presented approaches for quantifying the risk of incorrect resolution for a high-precision navigation system.
The parameter significance tests can also be conducted with the weighted squared norm of the estimators. The corresponding distributions under are given as follows (Teunissen, 2025), and the distributions under can be obtained with c = 0:
23
Based on the AF, AK, and AR estimators of c, four parameter significance tests can be defined, which are given in Table 1. The AK and AF significance tests can be formulated based on either the distributions of the corresponding estimators or their squared norms. For a given significance level α, the critical values for the AF and AK test can be obtained based on the CDF of the central chi-squared distribution with q degrees of freedom, denoted as , as shown in Table 1.
Owing to the multimodality of , two types of AR tests are defined with different acceptance regions: a region with high probability density (ARs) or a continuous ellipsoidal region around the origin, obtained with the squared norm of (ARn). An example of the two acceptance regions is shown in Figure 2. The critical values for the two AR tests cannot be obtained analytically, and the implementation of these tests will be discussed in the next sections.
Example of the AR acceptance regions for q = 1
Green and red rectangles indicate acceptance regions for an AR test with a high-probability-density acceptance region (ARs) and a continuous acceptance region (ARn), respectively.
3 IMPLEMENTATION AND POWER EVALUATION
3.1 AF and AK Significance Tests
In this section, we describe how to obtain the AF and AK critical values and how to compute the power corresponding to a certain size of model error c. Our description is based on tests with the squared norm of the estimators, with which the computation of critical values and powers can be directly carried out based on the CDF of the chi-squared distribution.
For a given significance level α, the AF and AK critical values are the 1 − α quantile of the CDF, denoted as . Under , the AF and AK test statistics follow non-central chi-squared distributions. The power of identifying a model error with size c can be computed based on the CDF of their distributions (Equation (23)) as follows:
24
Although the AF and AK tests have the same critical value, they have different powers because of the different non-centrality parameters under . The power of the AK test is no smaller than that of the AF test because the better precision of the AK estimator leads to a larger non-centrality parameter for the same c.
3.2 AR Significance Test with a High-Density Acceptance Region
3.2.1 Critical Value and Power Simulation
Owing to the multimodality of the distribution , both the critical value and the power of the ARs test can only be obtained through Monte Carlo simulation. This subsection describes the simulation procedure.
The simulation of the ARs critical value is summarized in the following steps, where is the number of samples to be used for simulating the critical value.
C1 Simulate the PMF of the integer estimator according to Section 2.2. The outcome is the probability mass for each z in a set of integers.
C2 Obtain samples of that follow the distribution of the AR estimator in Equation (22).
(a) For each integer z in the set of integers obtained in step C1, the number of samples is set to .
(b) Generate samples that follow a normal distribution , with c = 0.
(c) Repeat (b) for each integer vector with computed in (a) and collect all of the samples. This process produces samples of , denoted as .
C3 Compute the probability density for each sample according to Equation (22). Note that the infinite integer space in Equation (22) should be replaced by the finite integer set over which the PMF is obtained, as described in Section 2.2.
C4 Sort in ascending order; the -th ordered is taken as the critical value for the ARs test.
The power of the ARs test is a function of c. The steps for obtaining the power with samples are as follows.
P1 Follow steps C1–C3 to obtain samples of , denoted as . In this case, c in step C2(b) is not zero; instead, c is the size of the bias for which the test power will be evaluated.
P2 Compute the probability density for each sample according to Equation (22).
P3 Count the number of values smaller than simulated critical value , and denote this number as . The power for testing the model misspecification with size c is approximated as .
3.2.2 Number of Samples for the Critical Value
To conduct the ARs test, one must select the number of samples for simulating the PMF of the integer ambiguity estimator and the number of samples for the critical value . We conducted two experiments to provide insights into the number of samples that should be used. These experiments are based on a dual-frequency single-epoch DD Global Positioning System (GPS) observation model, and the ambiguity resolution success rate under is 91.7%.
For the experiment shown in the left graph of Figure 3, the number of samples is fixed at . The ARs critical value is obtained based on the ILS PMF simulated with three different numbers of samples and alternatively using the PMF of the IB estimator. For the experiment shown in the right graph of Figure 3, the number of samples for simulating the ILS PMF is fixed at . Three different values for the number of samples, , are used to simulate . The simulations with different setups were repeated 200 times in both experiments to evaluate the precision of the simulated . The experiments show the following:
(Left) Simulated ARs critical values for with samples of and different numbers of samples for the ILS PMF , as well as the critical value simulated with the IB PMF; (right) simulations with samples for the ILS PMF simulation and different numbers of samples for , i.e.,
The simulations in both graphs are repeated 200 times, shown along the horizontal axis. The empirical standard deviation of the simulated in each experiment is given in the legend.
The precision of the simulation is more sensitive to than to . As increases from 105 to 107, the simulation standard deviation is reduced by approximately 56%, whereas the same change in reduces the standard deviation by 90%.
In the first experiment, when is increased 10 fold from 106 to 107, the gain in simulation precision is marginal.
The critical value obtained with the IB PMF is not the same as that obtained with the simulated ILS PMF. Therefore, the IB PMF cannot be used as an approximation of the ILS PMF to simulate the ARs critical value.
Based on these analyses, is used in this contribution to simulate the ILS PMF, and is used to obtain the ARs critical value and the power. In practical applications of the ARs test, is sufficient to simulate the ILS PMF, as it provides a critical value simulation precision similar to that of .
3.3 AR Significance Test with an Ellipsoidal Acceptance Region
The implementation of the ARn test is also based on the PMF of the integer estimator, which should be computed or simulated as described in Section 2.2. Subsequently, the ARn test can be implemented and evaluated numerically without simulation because of its ellipsoidal acceptance region. The probability of the test statistic being inside the acceptance region is computed as follows:
25
with under and c = 0 under (Teunissen, 2025).
To obtain the ARn critical value, a root-finding method can be used to find the root , such that , with α being the specified significance level. The critical value of the AR and AK tests, , can be used as the initial value for the root-finding algorithm.
For a certain magnitude of model error c, the power of the ARn test, , can be computed from Equation (25) under .
4 NUMERICAL EXPERIMENT
We evaluate the performance of the parameter significance tests by computing their power as a function of the size of the parameter c. The power functions of five parameter significance tests will be evaluated and compared: the AK test, AF test, the AR test with an ellipsoidal acceptance region based on the ILS estimator, and AR tests with a high-probability-density acceptance region based on the ILS and IB estimators. Section 4.1 presents experiments for identifying a blunder in pseudorange and a phase cycle slip in the GNSS mixed-integer observation model. The power functions in this subsection are obtained over 20 values of bias c. Section 4.2 shows experiments for testing one-dimensional ionosphere delay and troposphere delay. The power functions in this subsection are obtained over 100 values of bias c to capture the sharp reductions in the power functions. Our numerical experiments are based on a single-constellation model with GPS satellites. The AR tests are expected to perform better when a stronger multi-constellation model is used, because the float ambiguity estimator is more precise.
4.1 Blunder in Pseudorange and Carrier-Phase Cycle Slip Test
In this subsection, we show the power functions of mixed-integer parameter significance tests to identify a blunder in pseudorange (Figure 4) and a phase cycle slip (Figure 5), together with the associated satellite skyplot and distributions of c estimators. Results for three experiments are shown in each figure for typical cases where the performance of the AR test is (i) poorer than that of the AF test, (ii) between that of the AR test and AK test, and (iii) very close to that of the AK test. It is demonstrated that the AR tests can outperform the AF test even if the ambiguity resolution success rate is not close to one. We also explain why the AR test may be poorer than the AF test and propose a criterion to predict cases in which the AR test performs better.
Skyplots, distributions of c estimators, and power functions for a blunder in a pseudorange test with a single-epoch dual-frequency double-differenced GPS observation model The top row shows results for an experiment in which the AR tests can perform worse than the AF significance test. In the middle row , the performance of the AR tests lies between that of the AF and AK tests. The bottom row ( ) shows that the results of tests can be close to those of the AK significance test.
Skyplots, distributions of c estimators, and power functions for the carrier-phase cycle slip test with a single-frequency double-differenced GPS observation model The top row shows results for an experiment in which the AR tests can perform worse than the AF significance test. In the middle row , the performance of the AR tests lies between that of the AF and AK tests. The bottom row ( ) shows that the results of the AR tests can be close to those of the AK significance test.
As shown in Figures 4 and 5, the distributions of the AR test statistics are close to unimodal for the blunder in the pseudorange test and the carrier-phase cycle slip test. Therefore, the acceptance regions of the ARs and ARn tests are the same. It follows that the power functions of the ARs and ARn tests overlap. According to the numerical experiments, the ARn test is recommended for a one-dimensional blunder in a pseudorange test and the cycle slip test, as it provides the same performance as the ARs test and is more computationally efficient.
For the experimental results depicted in the middle row of Figure 4 with and Figure 5 with , the ARs and ARn tests can provide higher detection power than the AF test, even if the ambiguity resolution success rate is not close to one. For the experiment represented in the bottom row of Figure 4, the AR tests can perform very close to the AK case with . The ARs test with the ILS estimator (ARs-ILS) performs better than that with the IB estimator (ARs-IB) because the ILS estimator has a higher success rate and a sharper PMF.
The top row in Figures 4 and 5 shows that the AR tests can perform poorer than the AF test, owing to the heavier tail of the distribution of compared with the AF estimator . Distributions of the AR (with the ILS estimator) and AF estimators of c are shown in Figure 6 on logarithmic axes. The left and right graphs present distributions corresponding to the top and bottom experiments in Figure 4, respectively. When the success rate is low and the shifts of the components of (see Equation (22)), , are large, the distribution of can have a heavier tail than the AF estimator . As a result, the AR acceptance region can be larger than the acceptance region of the AF test, as shown in the left graph of Figure 6, which leads to poorer performance. When the success rate is high or the shifts are small, the distribution of will be more concentrated, as shown in the right graph of Figure 6. The AR acceptance region is narrower, and the AR test performs better than the AF test.
Distributions of ARs (with ILS, blue) and AF estimators (orange)
The left and right graphs show results for the experiment represented in the top and bottom panels of Figure 4, respectively, on logarithmic vertical axes. Dashed lines indicate acceptance regions for the two tests with α = 0.01.
For the experiment corresponding to the top row in Figure 4 with and (blue curve), the AR test performance is poorer than that of the AF test. Meanwhile, the experiment represented in the middle row of the same figure shows that the AR test can perform better than the AF test, with a lower success rate of . Hence, we cannot determine whether the AR test performs better based on the success rate. To predict whether the AR test is better than the
AF test, we propose an indicator by comparing the corresponding estimators. For tests of a blunder in pseudorange and carrier-phase cycle slip, the distributions of the AR estimator are close to unimodal. The test will be better if the corresponding estimator is better, in the sense that its distribution is more concentrated. Based on this idea, we predict whether the AR test is better by comparing the probability of and lying in the AF acceptance region. For a certain significance level α, it is better to conduct the AR test if the following holds:
26
where is the acceptance region of the AF test and by definition.
The calculation of is described in the Appendix. For the six experiments shown in Figures 4 and 5, the probabilities are provided in Table 2. Each value in the table is associated with one power function of the AR significance test in the figures. For instance, for the AR test with the ILS estimator and in the experiment represented in the top row of Figure 4, the probability is smaller than , indicating that the AR test performs poorer than the AF test, which is consistent with the behavior of the associated (blue) power function shown in the top row in Figure 4. The probabilities that do not satisfy Equation (26) are marked in red, which correctly predict the experiments in which the AR test performs poorly. Table 2 shows that the criterion in Equation (26) can successfully predict whether the AR test achieves a better performance. It is recommended to that this criterion be applied before AR tests are performed for a blunder in pseudorange or cycle slip to avoid situations in which the AR test may perform poorly.
4.2 Ionosphere and Troposphere Delay Test
In the relative positioning scenario, one might face non-negligible differential atmosphere delays or atmosphere delays that remain after the application of network corrections. Figures 7 and 8 show the performance of the parameter significance tests for identifying one-dimensional ionosphere delay and troposphere delay, respectively. The top and middle rows in each figure show results for experiments with full ambiguity resolution, and results with partial ambiguity resolution (PAR) are shown in the bottom row. The choice of the subset of ambiguities to be resolved is driven by the resolution success rate. The subset is chosen to contain the maximum number of ambiguities and to have a resolution success rate larger than a minimum required value (Teunissen et al., 1999), which is set to 0.999 in our experiments.
The top and middle rows show skyplots, distributions of c estimators, and power functions for the ionosphere delay test with a dual-frequency double-differenced GPS observation model, with an ambiguity resolution success rate (SR) of and , respectively. The experiment represented in the bottom row follows the same setup as the experiment represented in the middle row, with PAR (12 out of 14 ambiguities are fixed). TECU: total electron content units
As shown in Figures 7 and 8, the ARs test can generally provide high power in identifying ionosphere delay or troposphere delay, even if the ambiguity resolution success rate is not close to one. The power can be very low for specific sizes of bias c, leading to spiky power functions. This behavior differs from the power functions for testing a blunder in a pseudorange or a carrier-phase cycle slip. This trend is driven by the strong multimodality of the AR estimator's distribution, , which can be well explained according to Section 4 of the work by Teunissen (2025). For the ionosphere delay and troposphere delay estimators, (components of ; see Equation (22)) is driven by the variance of the carrier-phase observable and is thus very sharp. Meanwhile, the shift of is inversely proportional to the phase variance and is thus significant. As a result, when the ambiguity resolution success rate is not close to one, we can expect for the ionosphere delay and troposphere delay estimators to have multiple well-separated sharp peaks. Figure 9 shows for and for the experiment represented in the top row of Figure 7 as an example. The power of the ARs test is equal to the integration of outside the high-density region of . Therefore, not all of the peaks contribute to the high-density ARs acceptance region; rather, only the high peaks that are above the ARs critical value contribute, which leads to sharp decreases in the power function. When the size of c is close to a high peak of , as shown in the left graph of Figure 9, the probability that lies outside the ARs acceptance region will be low, leading to a low power.
The top and middle rows show skyplots, distributions of c estimators, and power functions for the troposphere delay test with a dual-frequency double-differenced GPS observation model, with an ambiguity resolution success rate of and , respectively. The experiment represented in the bottom row follows the same setup as the experiment represented in the middle row, with PAR (9 out of 10 ambiguities are fixed).
Distributions of the AR estimator in the experiment represented in the top row of Figure 7, under (blue) and (orange). (a) , (b) TECU, (c) .
In panel (b), the high-density regions of are fully outside those of , leading to a very high power. In panel (c), the high-density regions of overlap with those of , which explains the first decrease in the ARn-ILS power function in the top row of Figure 7.
The ARs test with the IB estimator may perform similarly to that with the ILS estimator, as shown in the middle row of Figures 7 and 8. The distribution of the IB-resolved bias estimator almost overlaps with that of the ILS estimator, depending on the correlation between ambiguities after the decorrelation transformation (Teunissen, 1995). If the transformed ambiguities are nearly uncorrelated, the IB estimator can be similar to the ILS estimator (Teunissen, 2017). If this is not the case, as shown in the top row in Figure 8, the IB estimator can have a very spiky power function compared with the ILS estimator. The distribution of the IB-resolved bias estimator (red) differs from that of the ILS estimator (yellow). Because the ILS is a better estimator with a more concentrated distribution (PMF), the corresponding AR distribution has fewer high peaks. Moreover, a power function corresponding to a higher level of significance may contain fewer decreases, as shown in the middle row in Figure 7, because a higher significance level is associated with a larger ARs critical value . Hence, fewer peaks will contribute to the ARs acceptance region, leading to fewer drops in the power function.
The ARn test has a stepped power function for ionosphere delay and troposphere delay, which behaves very differently from the ARs test. This behavior is driven by the continuous ARn acceptance region and the multimodal distribution The ARn power function will increase significantly only when a peak of shifts outside the acceptance region with an increase in c, which lead to the “steps” in the ARn power function.
Although the ARs test can provide significantly more power than the AF test, sharp reductions in its power function can be observed. As shown in the top and middle rows of Figure 7, the ARs power can be lower than the AF power for several values of bias, which is undesirable. One idea for reducing the decreases in the ARs power function is to conduct the ARs test with partially resolved ambiguities (Teunissen, 2025). The multimodality of can be reduced by using the more peaked PMF of the partial integer ambiguity estimator, and the success rate of fixing a subset of ambiguities can be very close to one, e.g., larger than 0.999. In this case, the performance of the partial ARs (PARs) test is very close to the ambiguity partially-known test owing to the high PAR success rate, and the PARs power function is a smooth curve that is no poorer than that of the AF test. The bottom rows in Figures 7 and 8 show the distributions of the PAR estimators of c and the power functions of the PAR tests. Because of the very high success rate, the PAR estimator of c follows a unimodal distribution, and the PARs and PARn tests perform the same. We observe that the improvement in the power of the PAR test is generally less significant compared with the full ambiguity resolution case. This can be explained by Lemma 2 of Teunissen (2025):
27
where ADOP and a smaller ambiguity dilution of precision (ADOP) leads to a higher success rate. The above lemma indicates that when the increase in the ADOP from to is limited (the above ratio is close to 1), the improvement in the precision of the bias estimator will also be limited. In the PAR case, the success rate under is very high and satisfies . Therefore, the ratio will be limited, and the precision improvement in the bias estimator will be less significant.
The idea of combining the ARs test with the PARs test has been previously proposed (Teunissen, 2025), where is accepted only if it is accepted in both tests and the combined test provides a higher power than each of the individual tests. The motivation is to take the high power of the ARs test and reduce the decreases in its power function by combining it with the PARs test. Figure 10 compares the power functions of the AF, ARs, and PARs tests and the combined ARs + PARs test (purple). The power function of the combined test is always larger than that of each of the individual tests. Therefore, combining the PARs test with the ARs test can reduce the decreases in the ARs power to be no lower than that of the PARs test and guarantees that the combined test is always better than the AF test. We note that the significance level of the combined test also increases, approximately doubling that of the individual test. The significance level of the individual test can be set as half of the required value when the test combination is applied (Teunissen, 2025).
Power functions of the AF, ARs, and PARs tests with α = 1% and a combined ARs + PARs test
The left and right graphs are based on the experiment represented in the bottom row of Figure 7 for the ionosphere delay test and Figure 8 for the troposphere delay test, respectively. The significance level of the combined test is α = 1.9% for both experiments.
5 SUMMARY AND CONCLUSION
Based on the AR estimator of the bias, two AR-based parameter significance tests—namely, the ARs test and the ARn test—have been developed to validate the GNSS mixed-integer observation model (Teunissen, 2025). This contribution described how the ARs and ARn tests should be applied and evaluated their performance by computing their power as a function of bias magnitude.
We first reviewed the theory of the ARs and ARn tests, as well as significance tests based on the AF and AK bias estimators. We outlined the process for obtaining the PMF of both the ILS and IB estimators, which are essential for applying the AR-based tests. We introduced a method for obtaining the ARs critical value and power through Monte Carlo simulation. For the ARn test, the critical value can be obtained numerically through a root-finding method, and the power can be computed analytically. We then evaluated the performance of the AR test for identifying a blunder in pseudorange, a carrier-phase cycle slip, and ionosphere and troposphere delays, based on short-baseline double-differenced GNSS observation models.
The performance of the AR test is driven by the distribution of the AR bias estimator, denoted as . For a blunder in a pseudorange and the carrier-phase cycle slip test, is nearly unimodal. Therefore, the ARs and ARn tests have the same performance, because the high-probability-density acceptance region of the ARs test is also continuous. The test with the ILS estimator outperforms the test with the IB estimator, because the ILS estimator has a more concentrated PMF. Notably, the AR significance tests can perform better than the AF test, even if the success rate is not close to one. However, the AR power function can be below the AF power function. To predict when the AR test is better than the AF test, we propose an easy-to-compute criterion based on evaluating the probability that the AR bias estimator lies in the acceptance region of the AF test. The steps for obtaining this probability are described in the Appendix. Experiments confirm the effectiveness of this criterion, which is recommended to be applied prior to an AR test for a blunder in pseudorange or a cycle slip to avoid situations in which the AR test may perform poorly.
For the ionosphere and troposphere delay tests, the distribution of the AR bias estimator has a strong multimodality. In such cases, the ARs test behaves differently from the ARn test, because the ARs acceptance region is discontinuously distributed under the peaks of and the ARn acceptance region is continuous. Although the ARs test can achieve high power even at a relatively low ambiguity resolution success rate (e.g., ), its power may decrease sharply for specific bias magnitudes, resulting in spiky power functions. The ARn test produces a stepped power function. We explained the shapes of the ARn and ARs power functions based on the shape of and their acceptances regions. To mitigate the reductions in the ARs power function, the power functions of PAR tests were provided, in which a subset of ambiguities are resolved to ensure that the ambiguity resolution success rate is close to one. PARs and PARn tests show similar performance, and their power functions are smooth and never fall below that of the AF test. We also evaluated the combination of the ARs and PARs tests, which always performs better than each of the individual tests.
HOW TO CITE THIS ARTICLE:
Yin, C., Teunissen, P.J.G., & Tiberius, C.C.J.M. (2026).Ambiguity-resolved parameter significance test: implementation and performance for GNSS model validation.NAVIGATION, 73. https://doi.org/10.33012/navi.778
CONFLICT OF INTEREST
The authors declare no potential conflicts of interest.
APPENDIX
In this Appendix, we describe how to obtain the probability , with as the acceptance region of the AF significance test, which aids in determining whether the AR or AF significance test should be used.
If the test is multidimensional with can be approximated by the proportion of samples of the AR estimator lying in . The samples of under are obtained through the steps described in Section 3.2. Equation (25) cannot be used in the calculation because the shape of is determined by whereas the distribution of is a weighted sum of normal distributions with variance-covariance matrix .
If the test is one-dimensional with can be computed from the CDF of the normal distribution as follows:
Represent the AF acceptance region as , where is the standard deviation of
According to the distribution of in Equation (23), we have the following:
28
where is the variance of the AK estimator . The probability of the normal distribution with mean and variance inside can be computed as follows:
29
where is the CDF of the standard normal distribution.
Acknowledgments
Chengyu Yin's work in the “I-GNSS positioning for assisted and automated driving” project (18305) was funded by the Dutch Research Council (NWO). For this research, the Dutch national e-infrastructure Snellius supercomputer was used with the support of the SURF Cooperative under grant no. EINF-7002 and EINF-12666.
This is an open access article under the terms of the Creative Commons Attribution License, which permits use, distribution and reproduction in any medium, provided the original work is properly cited.
REFERENCES
- ↵Baarda, W. (1968). A testing procedure for use in geodetic networks (Publications on Geodesy, New Series, Vol. 2, No. 5). Netherlands Geodetic Commission.
- ↵Braasch, M. S. (2017). Multipath. In P. J. G. Teunissen & O. Montenbruck (Eds.), Springer Handbook of Global Navigation Satellite Systems (pp. 443–468). Springer.
- ↵De Jong, K., & Teunissen, P. J. G. (2000). Minimal detectable biases of GPS observations for a weighted ionosphere. Earth, Planets and Space, 52, 857–862. https://link.springer.com/article/10.1186/BF03352295
- ↵Drevelle, V., & Bonnifait, P. (2011). A set-membership approach for high integrity height-aided satellite positioning. GPS Solutions, 15, 357–368. https://link.springer.com/article/10.1007/s10291-010-0195-3
- ↵El-Mowafy, A. (2014). GNSS multi-frequency receiver single-satellite measurement validation method. GPS Solutions, 18, 553–561. https://link.springer.com/article/10.1007/s10291-013-0352-6
- ↵Eueler, H.-J., & Goad, C. C. (1991). On optimal filtering of GPS dual frequency observations without using orbit information. Bulletin Géodésique, 65, 130–143. https://link.springer.com/article/10.1007/BF00806368
- ↵Feng, S., Ochieng, W., Moore, T., Hill, C., & Hide, C. (2009). Carrier phase-based integrity monitoring for high-accuracy positioning. GPS Solutions, 13, 13–22. https://link.springer.com/article/10.1007/s10291-008-0093-0
- ↵Hernández-Pajares, M., Juan, J. M., Sanz, J., Aragón-Àngel, À., García-Rigo, A., Salazar, D., & Escudero, M. (2011). The ionosphere: Effects, GPS modeling and the benefits for space geodetic techniques. Journal of Geodesy, 85, 887–907. https://link.springer.com/article/10.1007/s00190-011-0508-5
- ↵Hewitson, S., & Wang, J. (2006). GNSS receiver autonomous integrity monitoring (RAIM) performance analysis. GPS Solutions, 10, 155–170. https://link.springer.com/article/10.1007/s10291-005-0016-2
- ↵Hofmann-Wellenhof, B., Lichtenegger, H., & Collins, J. (2012). Global Positioning System: Theory and Practice. Springer Science & Business Media. https://link.springer.com/book/10.1007/978-3-7091-6199-9
- ↵Khanafseh, S., & Pervan, B. (2010). New approach for calculating position domain integrity risk for cycle resolution in carrier phase navigation systems. IEEE Transactions on Aerospace and Electronic Systems, 46(1), 296–307. https://ieeexplore.ieee.org/document/5417163
- ↵Kok, J. J. (1984). On data snooping and multiple outlier testing (Vol. 30). US Department of Commerce, National Oceanic; Atmospheric Administration.
- ↵Lawrence, D., Langley, R. B., Kim, D., Chan, F.-C., & Pervan, B. (2006, April). Decorrelation of troposphere across short baselines. In Proceedings of IEEE/ION Position, Location, and Navigation Symposium (PLANS) (pp. 94–102). https://doi.org/10.1109/PLANS.2006.1650592
- ↵Leick, A., Rapoport, L., & Tatarnikov, D. (2015). GPS satellite surveying. John Wiley & Sons.
- ↵Massarweh, L., Verhagen, S., & Teunissen, P. J. G. (2025). New LAMBDA toolbox for mixed-integer models: Estimation and evaluation. GPS Solutions, 29(14), 1–9. https://doi.org/10.1007/s10291-024-01738-z
- ↵Odijk, D., & Teunissen, P. J. G. (2008). ADOP in closed form for a hierarchy of multi-frequency single-baseline GNSS models. Journal of Geodesy, 82, 473–492. https://doi.org/10.1007/s00190-007-0197-2
- ↵Odijk, D., Zhang, B., Khodabandeh, A., Odolinski, R., & Teunissen, P. J. G. (2016). On the estimability of parameters in undifferenced, uncombined GNSS network and PPP-RTK user models by means of S-system theory. Journal of Geodesy, 90(1), 15–44. https://doi.org/10.1007/s00190-015-0854-9
- ↵Odolinski, R., & Teunissen, P. J. G. (2022). Best integer equivariant position estimation for multi-GNSS RTK: A multivariate normal and t-distributed performance comparison. Journal of Geodesy, 96(3), 1–14. https://doi.org/10.1007/s00190-021-01591-9
- ↵Perfetti, N. (2006). Detection of station coordinate discontinuities within the Italian GPS Fiducial Network. Journal of Geodesy, 80, 381–396. https://doi.org/10.1007/s00190-006-0080-6
- ↵Teunissen, P. J. G. (1995). The least-squares ambiguity decorrelation adjustment: A method for fast GPS integer ambiguity estimation. Journal of Geodesy, 70, 65–82. https://doi.org/10.1007/BF00863419
- ↵Teunissen, P. J. G. (1998). The distribution of the GPS baseline in case of integer least-squares ambiguity estimation. Artificial Satellite, 33(2), 65–75.
- ↵Teunissen, P. J. G. (1999a). An optimality property of the integer least-squares estimator. Journal of Geodesy, 73, 587–593. https://doi.org/10.1007/s001900050269
- ↵Teunissen, P. J. G. (1999b). The probability distribution of the GPS baseline for a class of integer ambiguity estimators. Journal of Geodesy, 73, 275–284. https://doi.org/10.1007/s001900050244
- ↵Teunissen, P. J. G. (2001, June). GNSS ambiguity bootstrapping: Theory and application. In Proceedings of International Symposium on Kinematic Systems in Geodesy, Geomatics and Navigation (pp. 246–254).
- ↵Teunissen, P. J. G. (2005). On the computation of the best integer equivariant estimator. Artificial Satellites, 40(3), 161–171.
- ↵Teunissen, P. J. G. (2017). Carrier phase integer ambiguity resolution. In P. J. G. Teunissen & O. Montenbruck (Eds.), Springer Handbook of Global Navigation Satellite Systems (pp. 661–685). Springer.
- ↵Teunissen, P. J. G. (2018). Distributional theory for the DIA method. Journal of Geodesy, 92(1), 59–80. https://doi.org/10.1007/s00190-017-1045-7
- ↵Teunissen, P. J. G. (2024). The ambiguity-resolved detector: A detector for the mixed-integer GNSS model. Journal of Geodesy, 98(9), 1–16.
- ↵Teunissen, P. J. G. (2025). A significance test for mixed-integer models with application to GNSS. GPS Solutions, 29(4), 200. https://link.springer.com/article/10.1007/s10291-025-01952-3
- ↵Teunissen, P. J. G., Joosten, P., & Tiberius, C. C. J. M. (1999, January). Geometry-free ambiguity success rates in case of partial fixing. In Proceedings of the 1999 National Technical Meeting of the Institute of Navigation (pp. 201–207). https://www.ion.org/publications/pdf.cfm?articleID=669
- ↵Teunissen, P. J. G., & Khodabandeh, A. (2015). Review and principles of PPP-RTK methods. Journal of Geodesy, 89(3), 217–240. https://doi.org/10.1007/s00190-014-0771-3
- ↵Teunissen, P. J. G., & Kleusberg, A. (2012). GPS for Geodesy. Springer Science & Business Media.
- ↵Teunissen, P. J. G., & Montenbruck, O. (2017). Springer Handbook of Global Navigation Satellite Systems. Springer.
- ↵Verhagen, S. (2003). On the approximation of the integer least-squares success rate: Which lower or upper bound to use. Journal of Global Positioning Systems, 2(2), 117–124.
- ↵Verhagen, S., Li, B., & Teunissen, P. J. G. (2013). Ps-LAMBDA: Ambiguity success rate evaluation software for interferometric applications. Computers & Geosciences, 54, 361–376. https://doi.org/10.1016/j.cageo.2013.01.014
- ↵Wang, K., El-Mowafy, A., Qin, W., & Yang, X. (2021). Integrity monitoring of PPP-RTK positioning; part I: GNSS-based IM procedure. Remote Sensing, 14(1), 44. https://www.mdpi.com/2072-4292/14/1/44
- ↵Wanninger, L. (2004, September). Ionospheric disturbance indices for RTK and network RTK positioning. In Proceedings of the 17th International Technical Meeting of the Satellite Division of the Institute of Navigation (ION GNSS) (pp. 2849–2854). https://www.ion.org/publications/abstract.cfm?articleID=5968
- ↵Wu, S., Peck, S. R., Fries, R. M., & McGraw, G. A. (2008, May). Geometry extra-redundant almost fixed solutions: A high integrity approach for carrier phase ambiguity resolution for high accuracy relative navigation. In Proceedings of the IEEE/ION Position, Location and Navigation Symposium (PLANS) (pp. 568–582). https://doi.org/10.1109/PLANS.2008.4570100
- ↵Yang, L., Shen, Y., Li, B., & Rizos, C. (2021). Simplified algebraic estimation for the quality control of DIA estimator. Journal of Geodesy, 95(1), 14. https://link.springer.com/article/10.1007/s00190-020-01454-9
- ↵Yin, C., Teunissen, P. J. G., & Tiberius, C. C. J. M. (2025). Performance of ambiguity-resolved detector for GNSS mixed-integer model. GPS Solutions, 29(76), 1–17. https://link.springer.com/article/10.1007/s10291-024-01806-4
- ↵Zhang, W., Wang, J., El-Mowafy, A., & Rizos, C. (2023). Integrity monitoring scheme for undifferenced and uncombined multi-frequency multi-constellation PPP-RTK. GPS Solutions, 27(2), 68. https://link.springer.com/article/10.1007/s10291-022-01391-4















