Viyan

Viyan AI

Predicting Model Performance with CoRA-NAS

CoRA-NAS optimizes architecture selection by combining static proxy consensus with early-training performance extrapolation.

Neural architecture search typically relies on zero-cost proxies, metrics calculated on a single forward-backward pass, to predict which designs will perform best without training them to convergence. These proxies rely on structural assumptions, such as specific activation patterns or connectivity, that often fail when a model is moved to a new dataset. CoRA-NAS replaces this single-shot approach with a two-stage pipeline that corrects static estimates using early-stage training dynamics.

The Two-Stage Ranking Mechanism

The framework runs through two phases: CoRA-Rank and CoRA-Refine. CoRA-Rank creates a baseline by calculating disparate proxies, including synaptic saliency and Jacobian matrix variance. Because these metrics occupy different scales, the framework normalizes each into a rank-based distribution. It then computes an equal-weight consensus by averaging these normalized ranks. This process employs a target-free consensus gate to exclude proxies that exhibit low sensitivity, defined as those having a lower standard deviation across the population than the mean of the other proxies in the search space. By filtering these, the framework ensures only reliable indicators contribute to the final baseline.

CoRA-Refine uses a subset of architectures as anchors. It trains these anchors for a brief window to observe their early training trajectories. An ExtraTrees regression model takes the slope of the early training loss and the rate of accuracy improvement on a validation set to estimate the final performance of the architecture. Initial learning rate and loss reduction provide a strong, monotonic signal for final performance in well-behaved search spaces. The regression model predicts a residual correction value based on these early trajectories. The system applies this value to the initial static rank through a weighted summation, where the weight assigned to the residual is determined by the regressor's confidence in the trajectory. This update incurs approximately 1% of the cost of full training.

Metric Static Proxies CoRA-NAS
Search Cost Minimal ~1% of full training
Space Generalization Low High
NAS-Bench-101 Correlation Varies 0.715
NAS-Bench-201 Correlation Varies 0.946
Label Requirement None Validation Set

Refinement and Anchor Density

When applying this to a standard computer vision task like ResNet-based classification, the mechanism clarifies the distinction between early performance and total potential. A standard zero-cost proxy might suggest a high-performing architecture simply because it has many parameters or dense connections, but it cannot see how the weights actually update under gradient descent. By taking that 1% budget and running a few epochs, CoRA-NAS captures the actual curvature of the loss surface. The regression model recognizes that an architecture with a sharper, more consistent downward trend in loss is statistically more likely to reach higher accuracy than one that plateaus early or oscillates.

Determining the optimal number of anchors is an empirical exercise. If you sample too few architectures, the regression model suffers from high variance and fails to capture the nuances of the search space, leading to unreliable ranking adjustments. If you sample too many, the computation cost rises to meet that of standard training. The method provides a more accurate performance projection than purely structural proxies by grounding the estimates in observed training data. The primary risk is that the regression model is only as good as the correlation between the early trajectory and the final outcome; if an architecture has a high initial loss drop but suffers from later instability, the current estimation logic will overestimate its potential.

Sources