Changelog
Source:NEWS.md
ggRandomForests v4.0.0 (development)
?gg_partial_varpronow separates varPro versions in its missing-data and RMST-horizon notes. Before varPro 3.2.2,varpro()deletes incomplete cases silently and drops unrecognised arguments such asna.action, andpartialpro()drops a horizon passed through.... From 3.2.2 (kogalur/varPro#7),varpro()warns with the omitted count and records it inmodel.info$observations, and both functions stop with an error on an argument they do not recognise.partialpro()still has no horizon argument, soscale = "rmst"keeps supplying its own RMST(tau) learner. Documentation only; no function changed.gg_vimp()drops code that was meant to add arel_vimpcolumn but could never run: every fit, single-outcome included, takes the pivot branch, so no forest with stored importance ever returned it, and their output is unchanged. The@returnnow lists the columns actually returned (vars,set,vimp,positive), and theNAplaceholder for arandomForestfit without stored importance no longer carries an all-NArel_vimpcolumn.-
New
gg_ale_rfsrc()computes Accumulated Local Effects (Apley and Zhu,- for regression and classification forests, as a counterpart to
gg_partial_rfsrc(). Partial dependence averages the forest’s prediction over the joint distribution of the other predictors, which evaluates the forest at predictor combinations that never occur together when predictors are correlated. ALE only ever perturbs a predictor inside local neighbourhoods of its own observed values, so it does not extrapolate into those regions. Supplyingxvar2.namereturns the second-order (interaction) surface for a pair of continuous predictors, which is zero everywhere when the two act additively. Survival forests are not supported, the same limitationgg_shap()carries.
plot()andautoplot()methods draw the first-order curves as lines and bars faceted by variable, matchingplot.gg_partial_rfsrc()’s layout so the two can be read side by side, and the interaction surface as a heatmap. - for regression and classification forests, as a counterpart to
plot.gg_vimp(relative = TRUE)now plots relative VIMP: each variable’s VIMP divided by the largest VIMP in itsset, so the top variable reads 1 (per class for classification). The argument was documented but never read, so it silently plotted raw VIMP. It now defaults toFALSE. A set with no positive VIMP is scaled by its largest absolute VIMP, never divided by zero.plot.gg_vimp(nvar = )now keeps the topnvarvariables rather than the topnvarrows, so a classification plot no longer loses class panels.-
gg_partial_varpro()gainsscale = "prob_typical".partialpro()returns per-subject log-odds, and collapsing them to a curve takes an average and a back-transform; the ORDER is a modelling choice."prob"(unchanged, still the classification default) transforms per observation then averages, giving the mean predicted probability – the expected proportion of the cohort."prob_typical"averages on the log-odds scale then transforms once, giving the probability for a subject at the mean log-odds.They are different estimands and they disagree. The inverse logit is concave above zero and convex below it, so by Jensen
"prob"is pulled toward 0.5 at both ends, by more the more heterogeneous the cohort. Where the per-subject log-odds carry an SD near 4.5, a point reading 0.96 under"prob_typical"reads 0.74 under"prob"– large enough to change how a figure is read, so the choice should be deliberate. A figure captioned as a percentage of patients wants"prob".?gg_partial_varprosets out both.The distinction applies only to the
continuousframe; thecategoricalframe keeps values unaveraged, so both scales return the same numbers there. plot.gg_partial_varpro()gainsylim, pinning the shared y range across panels. It could not be set from outside: on thepanelsroute acoord_cartesian()added with&replaces the per-panel coordinate system and silently takes the per-panel x ranges with it (0-50 collapsed to the data’s 0-46), whilescale_y_continuous(limits = )is overridden by that same coordinate system.ylim = c(0, 1)now pins a probability axis so a flat curve reads as flat instead of filling the panel.plot.gg_partial_varpro()now defaultspaletteto"black". These figures are made for manuscripts, andlinetypeis mapped to the effect type as well, so the three estimators stay legible as solid, dotted and dashed with no colour at all. Pass any ColorBrewer name (palette = "Set1") for the colour scale, which separates two or three overlaid estimators faster on screen."mono"is a synonym for black and"grey"/"gray"give a flat grey. This changes rendered output: five vdiffr baselines were regenerated.plot.gg_partial_varpro()gainscomplement, plotting 1 - p and prefixing the y label with1 -. It reads a fit that targets one class as the probability of the other – a weaning-failure model shown as probability of successful weaning – without recomputingpartialpro()against the other target. Requires a probability scale (proborsurv); on the additive, multiplicative and unbounded scales 1 - x has no referent, so it errors rather than drawing something unreadable.plot.gg_partial_varpro()now warns, naming them, when arguments reach...that it does not use. Its own arguments sit after...and match by exact name, so a typo (or an argument from a newer version than the one installed) previously vanished without a word and left the default plot looking like a correct answer.plot.gg_partial_varpro()gains per-panel scale control.facet_wrap(scales = "free_x")gives each variable its own x range but a single shared x scale, so per-panel breaks, limits and axis titles were unreachable and a manuscript figure had to be hand-built oneggplot()per variable. A newpanelsdata frame – one row per panel, keyed byname, with optionalxlab,xmin,xmax,xbyandspancolumns – switches rendering to patchwork, where the x scale can vary between panels.panels = NULL(the default) is unchanged.plot.gg_partial_varpro()gainswhich, to return the continuous or categorical frame alone as a bareggplot. With both frames populated the method returns a patchwork, where+reaches only the last panel, so adding a scale or theme silently modified the categorical plot.whichis the supported way to get one plot to modify.plot.gg_partial_varpro()gainspoints,smooth,palette,ncol,point_size,point_alphaandlinewidth.palettetakes a ColorBrewer name and goes through ggplot2’s brewer scales, soRColorBrewerstays inSuggests. All default to the previous rendering.The survival vignette’s variable-dependence figure passed
time.labelswheregg_variable()readstime_labels. The dotted name matched nothing and was dropped without a warning, so the facet strips rendered as bare “1” and “3” instead of the intended “1 Year” and “3 Years”. Corrected; the figure now carries the labels its code always asked for.The survival vignette described
attr(gg_brier(rf), "crps_integrated")as a time-normalised score on the 0 to 0.25 Brier scale, then printed 1.44. The attribute isget.brier.survival()$crps, the raw area under the Brier curve in time units, and always has been. The vignette now says so and shows the normalised value (crps.std, and the right edge of the running CRPS curve); thegg_brier()help says the same. No change to any returned value.gg_brier()gains acrps_stdattribute,get.brier.survival()$crps.std: the integrated CRPS divided by the largest event time, so it reads on the Brier scale.print()andsummary()now report it as “CRPS (time-normalized)”, andsummary()labels the rawcrps_integratedas “integrated CRPS (time units)”. The value ofcrps_integratedis unchanged.The regression vignette’s partial dependence surface was flat in
rm. It overwrotenewx$rmand calledgg_partial_rfsrc()once perrmvalue, butnewxonly sets the evaluation grid andpartial.rfsrc()always averages over the training data, so all six curves were identical. It now holdsrmfixed throughxvar2.name, and the prose describes what the corrected figure shows: a modest interaction, not a strong one. Thenewxandxvar2.namedocumentation now says whatnewxdoes and does not control.Development line opened after the v3.2.0 CRAN release (forward-merged the v3.2.0 RMST/varPro fixes onto the dev line).
Begin the v4.0.0 development line: a Random Hazard Forests (RHF) visualization layer wrapping the ‘randomForestRHF’ package (added to Suggests). RHF support is gated — every gg_rhf* entry point checks
requireNamespace("randomForestRHF"). No change for users who do not install it.The consistency sweep distinguishes current CRAN software versions from supported minimum versions and standardizes the three package-qualified fit calls and object classes:
randomForestSRC::rfsrc()->rfsrc,randomForestRHF::rhf()->rhf, andvarPro::varpro()->varpro.Add a longitudinal RHF vignette covering
gg_rhf(),gg_auct(),gg_rhf_importance(), andgg_tune_rhf()from one saved analysis.gg_auct()/plot.gg_auct(): tidy wrapper and plot for time-varying AUC fromrandomForestRHF::auct.rhf()(RHF Phase 2). Returns a long frametime / auc / se / lower / upper / markerwith aniaucattribute (Uno + standardized integrated AUC);plot.gg_auct()draws AUC(t) with a bootstrap CI ribbon when available and a 0.5 reference line.gg_auct.rhf(object, marker, auct_fit = NULL)computesauct.rhf()internally or reuses a cached fit.gg_rhf_importance()/plot.gg_rhf_importance(): tidy wrapper and point matrix for time-localized variable priority fromrandomForestRHF::importance.rhf()(RHF Phase 3). It returnsvariable / time_window / time / time_index / start / stop / midpoint / n_risk / n_rules / priority, accepts a suppliedimportance_fitor calculates one when absent, and orders variables by their q90 priority over time windows. Priority is a ranking score, not a z-score; no selection cutoff is applied.gg_tune_rhf()/plot.gg_tune_rhf(): supplied-object-only inspection of atune.treesize.rhftree-size tuning path. The five returned columns aretreesize / metric / value / se / selected; the plot marks the selected size and draws an iAUC standard-error ribbon only when finite supplied iAUC standard errors are available.gg_tune_rhf()never recalculates tuning.Require
randomForestRHF (>= 2.0.3)in Suggests, and adopt its revised hazard semantics. From 2.0.0 the pointwise hazard is defined only where a grid point falls inside one of the case’s supplied(start, stop]intervals, and isNAin gaps and after the final stop. 2.0.0 left the cumulative hazard unmasked; 2.0.3 masks it as well, on its own rule, setting it toNAafter each case’s final stop while still holding it flat through an internal gap. The two masks therefore coincide on a fit whose cases carry a single interval each, and come apart only with time-dependent covariates.auct.rhf()can likewise return anNAAUC at the final grid time, where the censoring-weight denominator is undefined once the control set is nearly exhausted.gg_rhf()passes both masks through unchanged, sohazardandchfmay beNAwhere they previously were not;plot.gg_rhf()andplot.gg_auct()drop those cells before drawing, so a curve now ends with its case’s follow-up instead of reporting removed missing values on every plot. 2.0.0 also changes the default hazard aggregation (adaptive = TRUE), which shifts fitted values, and 2.0.3 corrects cumulative/dynamicauct.rhf(), which had inverted that curve; the RHF vdiffr baselines and the precomputed vignette analysis were regenerated against 2.0.3. This resolves issue #229, where the earlier reading (a small negative hazard, specific to the macOS arm64 binary) was wrong on both counts, and kogalur/randomForestRHF#1, the inverted cumulative/dynamic AUC.gg_auct()now errors rather than compute a cumulative/dynamic AUC it knows to be wrong.DESCRIPTIONasks forrandomForestRHF (>= 2.0.3), but R does not enforce aSuggestsversion at run time, so a session still carrying 2.0.0 previously got the inverted curve with no warning. The check is deliberately narrow: it applies only tomethod = "cumulative", since the incident definition never inherited the problem, and only whengg_auct()does the computation. A suppliedauct_fitis taken as given, because anauct.rhfobject records no version and may have been read from a file built elsewhere. The message names the installed version and points atmethod = "incident"as the alternative.gg_auct()gains amethodargument and now forwards...torandomForestRHF::auct.rhf().auct.rhf()defaultsmethodto"cumulative", andgg_auct()previously passed onlymarker, so the incident/dynamic definition could not be reached fromgg_auct()at all: the only route was to callauct.rhf()directly and hand the result back throughauct_fit.?gg_auctnow carries a note on choosing between the two definitions, which estimate different targets rather than better and worse versions of one. Forwarding...also makesbootstrap.repreachable, so the confidence ribbon no longer requires precomputing the fit.methodsits afterauct_fitin the signature, so positional calls are unchanged, and the default behavior is the same as before.plot.gg_partial_varpro(),plot.gg_partial(),plot.gg_vimp()andplot.gg_varpro()gain alabelsargument for human-readable variable names. It accepts a named character vector, a labelled data frame (readingattr(col, "label")), or a two-columnkey/labeldata frame. Variables with no label keep their raw name. Labels apply at draw time only — the returned object still carries raw variable names, so downstream consumers are unaffected. The argument reaches every branch of these methods, not just their default one:plot.gg_varpro()honours it on both the main panel and the class-conditional panel, andplot.gg_partial_varpro()honours it on survival path-C objects (those extracted withscale = "surv"or"chf"), which are handed off toplot.gg_partial_rfsrc().plot.gg_rhf_importance()also gainslabels, so the RHF priority matrix can carry human-readable variable names. It takes the same three shapes and falls back to the raw name per variable. The variable axis here isy, not a flippedx, so the labelled scale is the y scale. The q90 variable ordering and the raw names in the returned data are untouched, andautoplot.gg_rhf_importance()forwards the argument. Previouslylabelsfell through...intoggplot2::geom_point()and was dropped with only ggplot2’s generic “Ignoring unknown parameters” warning, so the call looked accepted and did nothing.plot.gg_beta_varpro(),plot.gg_ivarpro()andplot.gg_beta_uvarpro()gainlabelson the same terms, completing the varPro importance family. These three had been dropping the argument in complete silence: each declares...and does not use it, solabelswas absorbed with no warning, no error, and an unlabelled plot as the only symptom. Their facets are per class rather than per variable, so the class strips are left alone and only the variable axis is relabelled; in a faceted plot every panel is relabelled.plot.gg_shap()and the three exported mode functions it dispatches to,shap_importance(),shap_beeswarm()andshap_dependence(), gainlabels. Each of the three puts variable names somewhere different, so each honours the argument differently:shap_importance()labels a flipped discretexscale,shap_beeswarm()labelsydirectly because it does not flip, andshap_dependence()has no variable scale at all and substitutes the label into both axis titles instead. In that last modexvarstill matches on raw variable names, so the label is display only and passing a label where a variable name belongs is still an error. As with the varPro methods above,labelswas previously accepted and discarded in silence.plot.gg_variable()no longer declarestimeandtime_labels, two formals its body never read. They are parameters ofgg_variable(), the extractor, which bakes the horizon into the object beforeplot()runs, so the man page had been promising a horizon selection the method never performed. Supplying either now warns and names the call that works. ⚠️time_unitsmoved to after...in the signature as part of this: R partial-matches argument names only before..., so withtime_unitsahead of it a caller writingtime = 1191bound silently totime_unitsand died on its type check. Past the dots, matching is exact.time_unitsalways had to be named, so no working call changes.-
plot.gg_variable()sanity-checkstime_unitsagainst the data it describes. A year-like unit supplied against values above 150 warns, because that is almost always a forest fit on a smaller unit and produces an axis title wrong by a factor of- It warns rather than errors, and only in that one direction: small values labelled
"days"is ordinary, so there is no signal to check. The package still cannot derive the unit and does not try to. Scoped toplot.gg_variable()for now:plot.gg_rfsrc()also takestime_units, but agg_rfsrcobject has notimecolumn (its time points live invariable, which holds class names for a classification fit), so there is no unambiguous column to check and extending it is not a one-line change.
- It warns rather than errors, and only in that one direction: small values labelled
plot.gg_variable(),plot.gg_udependent()andplot.gg_sdependent()gainlabels, which completes the argument across every plot method in the package that renders variable names.plot.gg_variable()labels the facet strips in its panel plot, through all three faceting branches, and the x axis title in its individual plot; thetimefacet is untouched, because it facets by time rather than by variable, and the multi-time survival panel scopes its labeller to the variable dimension so a label key that collides with a time value cannot reach the time strips.plot.gg_udependent()labels the node text of its dependency network. There the display string is written to a separate vertex attribute and the igraphnameis left alone, becausenameis the key the edge-weight backfill matches on and rewriting it would break edge weights on graphs saved before those weights were stored.plot.gg_sdependent()is the plain case, a flipped discrete axis likeplot.gg_vimp().plot.gg_vimp():lblsis deprecated in favour oflabelsand will be removed in a future release. Its oldlength(lbls) >= length(vars)gate is also gone, so a partial label set is now honoured, falling back to the raw name per variable. Previously supplying fewer labels than variables silently applied none.gg_partial_varpro()orders variables by varPro importance (varPro::get.topvars()) whenobjectis supplied, andnameis now a factor, so facets follow importance order instead of being re-sorted alphabetically. Variables absent from the ranking keep their incoming order and are appended after the ranked block; none are dropped.gg_partial_varpro():nvarsnow selects the top n by importance. It previously took the first n elements of the partial-dependence list before any ranking was applied, returning an arbitrary subset with no symptom in the output.gg_partial_varpro()warns whenscale = "auto"cannot be resolved because noobjectwas supplied, instead of silently falling back to the generic “Partial Effect” axis. The fallback label itself was never wrong — it was honest about an unknown scale — but the silence around it was, so the fallback now says so.The
labelslookup now drops entries whose label or name is blank orNA, so a variable given an empty label falls back to its raw name rather than drawing blank axis or strip text. All three accepted input shapes now agree on the same information; previously the labelled-data-frame arm dropped blanks while a named vector orkey/labelframe kept them.gg_partial_varpro()now rejects an unnamedpart_dtawith a clear error instead of accepting it. The names of that list are the variable identities; without them the constructor cannot build anamecolumn at all, and the omission used to surface two calls later as an opaquefacet_wrap()failure about a missing faceting variable. An emptypart_dtaremains legal.The package’s own vignettes (
ggRandomForests-regression.qmd,ggRandomForests-survival.qmd) are moved off the now-deprecatedlblsargument ontolabels, so the shipped examples model the current API rather than the one being phased out.plot.gg_variable()andplot.gg_rfsrc()no longer hard-code a “year” time unit in survival axis titles. The unit was never derived from the data, so a fit measured in days –randomForestSRC::pbc, this package’s own canonical survival example, among them – rendered “Survival at 1191 year” for a horizon of 1191 days. The titles now read “Survival at 1191” and “time” by default, and a newtime_unitsargument on both methods restores an explicit unit:time_units = "days"gives “Survival at 1191 days” and “time (days)”. Users whose data really is in years should passtime_units = "years"to keep the word.plot.gg_variable()’s survival branch has visual regression cover for the first time. The branch forks onpaneland on whether the object carries one time or several, and the four resulting paths differ in both faceting and y-axis title; each now has avdiffrbaseline. The two single-time paths are the ones that render a time unit into the axis title, so a change to that title now surfaces as an SVG diff rather than resting on anexpect_equal()ofp$labels$y, which cannot see the rest of the panel. Tests only.
ggRandomForests v3.5.2
CRAN release: 2026-08-21
- Three help pages no longer render a stray backslash where a percent sign belongs. roxygen2 escapes
%for you, so the\%written in the roxygen prose ofcalc_auc(),gg_isopro()andplot.gg_isopro()reached the.Rdas\\%and rendered as50\%rather than50%. Documentation only. -
R CMD checkis back inside CRAN’s ten-minute budget. On the 3.5.1 win-builder run the vignette rebuild was 287s and the tests 195s of a 12-minute total, so both were cut at the source rather than moved around. Therfsrcvignettes grow smaller forests (Boston and iris at 100 trees, thepbcimpute-and-fit pair at 50 and 100) and coarser partial-dependence surface grids (6 and 5 points, from 10 and 8); the SHAP sections explain 25 rows against 30 background draws instead of 40 against 50. In the examples,gg_error()andplot.gg_error()grow 100 trees instead of 250, andgg_vimp()andplot.gg_vimp()50 instead of 100. The four heaviest test files,gg_udependent,gg_varpro,gg_variableandgg_vimp, nowskip_on_cran(); they still run in full underdevtools::test(). No function, argument or returned object changed. -
?gg_partial_varpronow documents varPro’s missing-data contract, which governs every fit this package plots. varPro has no imputation: each entry point grows a stump throughrandomForestSRC::rfsrcand inherits itsna.action = "na.omit", so any case missing a predictor or the outcome is deleted before the fit, silently.na.action = "na.impute"passed tovarpro()lands in...and is discarded without remark. The loss compounds as0.95^p, and the fitted object keeps only the post-deletion count, so neither the user nor this package can recover the original from the object – the check has to happen before the fit. - The same section covers imputing beforehand without inventing outcomes.
roughfix()andrandomForestSRC::impute()both fill every column handed to them, the outcome included, so a frame with missing outcomes comes back with manufactured responses and the release rules are fit partly to them. Documents von Hippel’s impute-then-delete, and the two cautions that follow: outcome-informed imputation crosses fold boundaries incv.varpro(), and a completed frame carries no imputation uncertainty into the curves.
ggRandomForests v3.5.1
- Test-only fix for the
gcc-UBSANadditional issue reported against 3.5.0. One test grew an isolation forest with no outcome, which makesrandomForestSRChandyvar.wt = numeric(0)to its native code and decrement that zero-length pointer (entry.c:184). The route was indirect:gg_partial_varpro()callsvarPro::partialpro(), which grows its ownisopro()forest and letsmethoddefault to"unsupv". The other live-partialpro()tests were alreadyskip_on_cran()’d, which left this one as the only one running on CRAN. It now requestsmethod = "rnd"itself; it asserts the same warnings over the same number of rows. No user-facing code changed –ggRandomForestsis pure R, and the undefined behaviour is upstream. - The comments in the varPro test fixtures claimed the report fires only for
isopro(method = "unsupv"). That was true of direct calls and wrong as a rule about the package, which is why thepartialpro()route went unnoticed. They now state the actual condition – anyrfsrcgrow reached without a formula – and name the paths that satisfy it. - An audit of the rest of the package found one more call on the same path: the
\donttestexample in?gg_partial_varproused the object path withoutmethod = "rnd". It did not fire on CRAN only because that check flavor did not run\donttestcode, which is CRAN’s setting to change rather than ours, so the example now passesmethod = "rnd"too. Every other varPro entry point is clean:gg_varpro(),gg_ivarpro(),gg_udependent(),uvarpro()(defaults to a formula-basedmethod = "auto") and everyisopro()call in the package, all of which namemethod = "rnd". -
?gg_partial_varpronow documents the underlying issue rather than leaving it to the tests. Anygg_partial_varpro(object = )call reaches it, becausepartialpro()grows its isolation forest withisopro()’s defaultmethod = "unsupv". It is benign – the pointer is formed, never dereferenced – and the fix belongs upstream (kogalur/randomForestSRCPR #478);method = "rnd"avoids it in the meantime. -
gg_roc()on anrfsrcforest now honorswhich_outcome = 0. The help page has always documented0as the numeric spelling of"all", but only the string was normalized, so0fell through topredicted[, 0]. That is a legal zero-column subset rather than an error, so the threshold sweep ran on empty input and returned a two-row frame with nosens/speccolumns, which then brokecalc_auc(). Both spellings now take the same route: a warning, and a fallback to class 1. The macro-average that will replace the fallback is still tracked under #72. - The three ROC entry points still disagree about what “all classes” means –
gg_roc()on arandomForestfit macro-averages,gg_roc()on anrfsrcfit falls back to class 1, and a directplot.gg_roc()call on a raw multi-class forest overlays one curve per class. That divergence is unchanged here, but?gg_rocand?plot.gg_rocnow say so instead of implying the paths agree. Both also correct a longer-standing claim: a raw forest passed to plainplot()never reachesplot.gg_roc()at all, becauserandomForestSRCandrandomForestregister their ownplotmethods and S3 dispatch prefers them. That branch is reachable only by naming the method outright.?gg_rocfurther stops advertising character class names on therfsrcpath, which only therandomForestmethod accepts. -
gg_partial_rfsrc()validatesrf_modelbefore using it. It read$xvarand$xvar.namesfirst, so a non-forest failed with base R’s “argument is of length zero” rather than naming the problem. It now matches the error style already used bygg_error(),gg_vimp(),gg_variable()andgg_rfsrc(). - The
pbcexamples on?gg_error,?plot.gg_error,?gg_vimpand?plot.gg_rfsrclost their editorial asides and a stray trailing comma in thedata()call. The munging block that all four repeated is now a single sharedinst/examples/pbc-setup.R, pulled in with@example, so the four pages cannot drift apart. - Every
rfsrcfit in an example now names an explicitntree. The examples had been takingrfsrc()’s 500-tree default, which is far more forest than an illustration needs; bounding them took the localR CMD checktotal from 4m44s to 3m16s. -
tests/testthat/test_lint.Rruns again, wrapped inskip_on_cran(). It had been commented out entirely, so the suite enforced nothing about style locally even though CI kept its own lint job. The guard keeps it off theR CMD checkclock.
ggRandomForests v3.5.0
CRAN release: 2026-08-04
-
plot.gg_varpro()no longer draws a phantom “NA” category whennvaris smaller than the number of variables the fit reports.$imp/$statsare truncated tonvar, but the per-tree overlay ($imp.tree) and the class-conditional data ($conditional) still carry every variable; re-levelling those to the truncated$implevels orphaned the extras toNA, which rendered as an empty box/bar. Those rows are now dropped, so only the displayed variables appear. - The vignettes now render their figures with
raggand quantise them to a 256-color palette, cutting the source tarball from 4.7 MB to 2.3 MB. The vignettes had never chosen a graphics device, so they fell through to the defaultpng(), which writes RGBA truecolor: an alpha channel these opaque plots never use, over tens of thousands of anti-aliased colors that PNG cannot compress. Figures are visually unchanged (mean pixel difference 1.55 on a 0-255 scale). Both steps are build-time only and degrade to no-ops whenraggormagickis absent, so a vignette rebuild without them still succeeds – at the old file size. - The varPro vignette now documents which variables a
varprofit actually makes available. A fit narrows the predictors twice –object$xvar.namesholds whatvarPro::partialpro()can reach,varPro::get.topvars()only the reported ranking – andpartialpro()silently drops any requested name outside the first set. The new section covers namingxvar.namesto get past the reported ranking,split.weight = FALSEto widen the candidate set itself, and the two arguments (nvar,sparse) that look like they should help and don’t. -
gg_partial_varpro()now warns when a name passed inxvar.namesis one thevarprofit cannot reach, instead of letting it disappear.partialpro()keeps only the names it finds inobject$xvar.namesand says nothing about the rest, so a request for twelve variables could come back with ten. The check runs beforepartialpro()does, so the warning arrives ahead of the computation rather than after it; it names every dropped variable and points atsplit.weight = FALSE. Supplyingpart_dtayourself is unchanged – the variables are already gone by then. The function’s examples now cover the object-driven path, which had none. -
gg_partial_varpro(scale = "chf")now computes the variables you name inxvar.namesinstead of every variable the fit can reach. Thechfpath routes throughgg_partial_rfsrc()rather thanpartialpro(), and it had never been given the variable list – so asking for one variable quietly did the work for all fourteen. This is the mirror image of thepartialpro()bug above: that one returns fewer variables than you asked for, this one returned all of them. Naming a variable the forest does not carry has always been an error and still is. Thepartialpro-only arguments (cut,nsmp) mean nothing on this path and are now ignored with a warning rather than in silence. - Fix:
gg_vimp()on arandomForestfit grown withimportance = TRUEnow reports the permutation importance you asked for. It was reporting node purity instead, and silently:randomForeststores%IncMSEandIncNodePurityside by side, andgg_vimp()stacked both into onevimpcolumn and ranked them together. The two are not commensurable – node purity runs in the thousands where%IncMSEruns in the tens – so every impurity row outranked every permutation row, and the truncation tonvarcut the permutation values away entirely. OnrandomForest(medv ~ ., Boston, importance = TRUE)the plot showedlstat = 12576.7(node purity) where the permutation value islstat = 62.4. Node purity is now left out of the ranking; readrandomForest::importance(object)if you want both. Fits grown withoutimportance = TRUEare unaffected – they only ever stored node purity, and that is still what you get. - Fix:
gg_vimp()on arandomForestclassification fit grown withimportance = TRUEnow reports permutation importance as well. That matrix mixes the same two scales – a permutation column per class plusMeanDecreaseAccuracy, alongsideMeanDecreaseGini– but it is wider than the single-outcome branch that picks one measure, so it skipped that branch and every column was ranked together.MeanDecreaseGinicame out the sole survivor: onrandomForest(Species ~ ., iris, importance = TRUE),gg_vimp()returned 4 rows of node purity where 16 rows of permutation importance were there to report. The per-class columns andMeanDecreaseAccuracyare all permutation measures on one scale, so they are now kept together and named in thesetcolumn, the way anrfsrcfit’sall/<class>columns already were; onlyMeanDecreaseGiniis dropped. - Fix:
which.outcomenow selects the column you asked for on arandomForestclassification fit.which.outcome = 0documented itself as overall importance and took column 1 to get it, andwhich.outcome = ktook columnk + 1for classk. Both are right for anrfsrcfit, whose$importanceleads with anallcolumn, and neither is right here: arandomForestmatrix opens on the classes and keeps the overall permutation measure inMeanDecreaseAccuracy, near the end. So0returned the first class labeled as overall – onrandomForest(Species ~ ., iris, importance = TRUE)it handed back setosa’s values, rankingPetal.WidthabovePetal.Lengthwhere the overall measure has them the other way round – and every class index was shifted by one,1giving versicolor. The columns are now resolved by name:0reachesMeanDecreaseAccuracy,kreaches classk, andwhich.outcome = 1agrees withwhich.outcome = "setosa". Fits grown withimportance = FALSEkeep noMeanDecreaseAccuracycolumn and their single measure answers to0as before. -
which.outcomenow names the measure it selected in thesetcolumn, for bothrfsrcandrandomForestfits. Asking for one measure reportedsetas the literal"vimp"– the pivot takessetfrom the source column name, and the selected column was named after thevimpcolumn it was about to be written into rather than after the measure it held. So the one path where you have to say which measure you want was the one path that would not tell you which measure you got.gg_vimp(rfsrc_iris, which.outcome = 0)now reportsset == "all",gg_vimp(rf_iris, which.outcome = 0)reportsset == "MeanDecreaseAccuracy", and both agree with the names the unfiltered pivot has always used. Values and ordering are unchanged, and plots are unaffected:plot.gg_vimp()only facets onsetwhen there is more than one of them, and selecting a measure leaves exactly one. -
nvarcounts variables again forrandomForestfits, not rows. It was applied after the multiclass pivot, where a frame holds one row per variable per measure, so it lopped whole measures off the end of the ranking instead of trimming the ranking itself. -
gg_vimp()now says in?gg_vimpthat arandomForestfit withoutimportance = TRUEstores onlyIncNodePurity, so the ranking is node purity rather than permutation VIMP, and nothing in the plot marks the difference. The example now passesimportance = TRUE. -
gg_error()now explains that the error trajectory israndomForestSRC’s to record, not ours:rfsrc()’sblock.sizedefaults toNULLunless you request importance, which stores the error at the final tree only, so a default fit givesgg_error()a single point rather than a curve –tree.err = TRUEalone does not change that. Grow withblock.size = 1for an error at every tree. The examples do this now; they had all been plotting one dot. -
gg_beta_varpro(): theimpcolumn is documented as the absolute coefficient.varPro::beta.varpro()wraps every coefficient it returns inabs(), so the sign is discarded upstream and never reaches us – the docs had said “Sign is real (direction of local association)”, which cannot be read off this output. Usegg_ivarpro()for a signed local estimator. -
gg_isopro(): the “What’s in the output” section now says the polarity flip is ours.varPro::isopro()’showbadis lower = more anomalous; we return1 - howbadso that higher = more anomalous. The section had credited that to the fit, contradicting this function’s own@return. - Added
gg_shap()andplot.gg_shap()(withshap_importance(),shap_beeswarm(),shap_dependence()) for SHAP explanations of regression and classification forests, wrappingkernelshap(Suggests). -
gg_shap()now enforces the integer contract onbg_nandwhich.classinstead of silently coercing them. Both are documented as integers, but were only loosely checked:bg_n = 1.9was truncated to 1 andbg_n = Inf(or any value above.Machine$integer.max) becameNA, whilewhich.class = 2.9passed the range check and then indexed column 2 – returning SHAP values for a class the caller never asked for. Non-whole, non-finite, out-of-range and non-scalar values now raise a clear error. Valid input is unaffected. - Added
print.gg_shap()andsummary.gg_shap().gg_shapwas the onlygg_*class without them, so it dumped every row at the REPL instead of showing a header.print()now gives the standard one-line header (with the variable and observation counts) andsummary()returns asummary.ggobject reporting the baseline, background-sample size, the explained class for classification fits, and the top variables by mean |SHAP|. - The package help page (
?ggRandomForests) now describes the whole current surface – the SHAP, Brier, varPro and unsupervised-varPro families were missing – and no longer claims thatplot()methods may return a list ofggplot2objects; each returns a single plottable object (aggplot, or apatchworkcomposite for the multi-panel methods). -
gg_partial()no longer lets survival partial dependence be mistaken for a probability.randomForestSRC::plot.variable()defaults tosurv.type = "mort", soyhatis mortality – an expected event count, not a value on [0, 1] – and it only superficially resembles a percentage.yhatis passed through unscaled (rescaling it would corrupt the quantity); instead the label describing what was plotted is carried on the object asattr(x, "ylabel")and used as the y-axis title byplot.gg_partial(). Note thatgg_partial_rfsrc()defaults topartial.type = "surv"and so does report survival probabilities: the two entry points report different quantities by default. (#15)
ggRandomForests v3.4.1
- The remaining
rfsrc/randomForestwrappers –gg_error(),gg_vimp(),gg_variable(),gg_rfsrc(), andgg_brier()– now havedefaultS3 methods, so a wrong-class input gives a clear “expected an ‘rfsrc’ or ‘randomForest’ object” error (naming the class it got) instead of R’s generic “no applicable method”. This finishes the dispatch-consistency pass started for the varPro family in 3.4.0. (gg_roc()keeps its existinggg_roc.rfsrcdefault, which accepts rfsrc-shaped objects.)
ggRandomForests v3.4.0
CRAN release: 2026-07-02
-
gg_isopro(),gg_beta_varpro(), andgg_ivarpro()now havedefaultS3 methods, so a wrong-class input gives a clear “expected a ‘’ object” error (naming the class it got) instead of R’s generic “no applicable method”. This makes the varPro-family wrappers consistent with gg_beta_uvarpro()/gg_sdependent(); the previously-unreachable inner class checks were removed. - Fix:
gg_partial_rfsrc()now computes partial dependence correctly forfactorpredictors. It was passing factor labels aspartial.valuestorandomForestSRC::partial.rfsrc(), which imposes a level by its integer code (internallyas.numeric(partial.values)). Character labels (“No”/“Yes”) becameNAand numeric-looking labels (“4”/“6”/“8”) became out-of-range codes, so every level collapsed to a single value (a flat categorical partial plot). The wrapper now passes the integer codes and relabels the output, matchingplot.variable(partial = TRUE)and the ground-truth partial dependence. The categoricalxis now returned as afactorin the model’s level order, so the plot keeps that order instead of re-sorting alphabetically. Continuous and numeric low-cardinality predictors are unaffected. -
gg_beta_uvarpro()/plot.gg_beta_uvarpro(): tidy wrapper and bar chart forvarPro::get.beta.entropy()– the unsupervised analogue ofgg_beta_varpro(). From auvarpro()fit it aggregates the per-region lasso coefficients intobeta_mean = colMeans(|beta|)per variable (most-important first), flags variables above a selection cutoff, and accepts a precomputedbeta_fitmatrix.print/summary/autoplotcompanions follow thegg_*conventions. -
gg_sdependent()/plot.gg_sdependent(): tidy wrapper and ranked lollipop forvarPro::sdependent()signal-variable detection. Returns one row per candidate variable (imp_score, graphdegree,signalflag) ranked byimp_score. Complementsgg_udependent()(the dependency graph) with the “which variables are signal” ranking; shares thebeta_fitentropy matrix. Follows theget.beta.entropy+sdependentworkflow from thevarPro::uvarpro()help (iowa-housing example). - New
uvarprovignette: a short, focused walk-through of the unsupervised varPro wrappers (gg_udependent(),gg_beta_uvarpro(),gg_sdependent()) on a singleuvarpro()fit, using the sharedbeta_fitmatrix. The three unsupervised sections were lifted out of thevarprovignette, which now points to the new one and covers the five supervised wrappers. - Fixed the main vignette’s
\VignetteIndexEntry, which still carried the template placeholder “Vignette’s Title” – it now reads “Exploring Random Forests with ggRandomForests” (the index entry CRAN lists, not the document title, was the stale one).
ggRandomForests v3.3.0
-
gg_partial_varpro(): classification partial plots now default to probability.scale = "auto"on a classification fit resolves to"prob"(P(Y = target class)) instead of raw log-odds;"odds"and"logodds"are options. The back-transform is applied before averaging (mean predicted probability). Thecausalcontrast is shown only on"logodds". -
gg_partial_varpro(): survival partial plots now default to survival probability.scale = "auto"on a survival fit resolves to"surv"(S(tau | x), bounded 0-1) via a new partialpro learner, instead of the unbounded ensemble-mortality score (still available viascale = "mortality")."surv"and"rmst"defaulttauto the median follow-up time whentimeis omitted – a units-safe, data-driven horizon (v3.2.0’srmstrequiredtime; this is a loosening). The resolvedtauis reported in a message and the axis label. -
plot.gg_partial_varpro(): documents what thecausal(virtual-twins) estimator is and when to use it, and explains why it is hidden on the bounded probability scales. - Documentation:
plot.gg_partial_varpro()gains a “Reading an RMST curve” section explaining how to interpret thescale = "rmst"y-axis – RMST(tau) is the expected event-free time within the first tau time-units (area under S(t) out to tau), read in the model’s own time units, bounded by tau, and higher-is-better (the opposite direction from ensemble mortality). It also notes that tau must be supplied in the fit’s time units, since a tau beyond the largest event time truncates to the full restricted mean. No code change.
ggRandomForests v3.2.0
CRAN release: 2026-06-23
- Fix (#118):
gg_varpro()no longer fails with the cryptic “arguments imply differing number of rows:, 0” when
varPro::importance()returns a degenerate importance table (0 rows, orpvariables with no usablezcolumn) – observed intermittently on survival fits where the release-rule step selects no variables. It now stops with a clear, specific message explaining the empty importance and suggesting a largerntree. The guard is scoped to the degenerate case only; well-formed fits (survival included) are unaffected – this is not a blanket survival-family block (cf. the reverted #116). - Fix:
gg_partial_varpro(scale = "rmst", time = tau)now drives the survival partial computation instead of only relabeling the y-axis.varPro::partialpro()has no time argument, so its default survival learner returns ensemble mortality at every horizon – multi-horizon RMST plots built that way differed only by Monte-Carlo noise, not bytau.scale = "rmst"now passespartialpro()an RMST(tau) learner that integrates the survival curve (integral_0^tau S(t) dt) fromobject$rf, so the curve genuinely depends ontau. This path recomputes fromobject(a survival fit) withpart_dta = NULL; a precomputedpart_dtacan only be relabeled, and the function now warns when you try. Also warns whentauexceeds the model’s event-time range (RMST is truncated there) and whentimeis passed to a scale that ignores it.Importsnow requiresvarPro (>= 3.1.0)(the version exposing thepartialpro()learnerargument this path relies on). - Fix:
gg_partial_varpro(scale = "surv"/"chf", model = ...)no longer errors when a variable yields an empty continuous or categorical frame (the survival path-Cmodel-label assignment now guards against a 0-row data.frame). -
gg_partial_varpro()(and thegg_partialpro()alias) now forward...tovarPro::partialpro()on the object-driven path. This restores control over which variables are computed (xvar.names,nvar) and the UVT step (cut,nsmp, …) for the RMST path, which must recompute fromobjectand so cannot accept a precomputedpart_dta. Without an explicitxvar.names,partialpro()falls back tovarPro::get.topvars(object), which can return few or no variables for some fits.
ggRandomForests v3.1.2
CRAN release: 2026-06-13
- CRAN fix: skip only the single test grow that trips the upstream
randomForestSRCgcc-UBSAN report atentry.c:184— the unsupervised isolation forest ingg_isopro(varPro::isopro(method = "unsupv")). Only an unsupervised grow has a 0-lengthyvar.wt, the vectorrfsrcGrowdecrements to an out-of-bounds pointer; supervised grows are unaffected. We verified this under-fsanitize=undefined: of every varPro/rfsrc grow in the test suite, onlyisopro(method = "unsupv")firesentry.c:184.make_iso_fit()therefore callsskip_on_cran()only formethod = "unsupv". ggRandomForests is pure R and unchanged. - The broader
skip_on_cran()guards added in v3.1.1 (thevarpro,uvarpro,ivarpro,beta.varpro, andisopro(method = "rnd")test fixtures) are removed: those grows are supervised (or synthetic-supervised) and gcc-UBSAN-clean, so they run on CRAN again, restoring that test coverage. The upstream issue is fixed inrandomForestSRCand pending a CRAN release.
ggRandomForests v3.1.1
- CRAN fix: the varPro tests now call
skip_on_cran()so they do not run on CRAN’s check machines, including the gcc-UBSAN additional check. They were triggering an upstreamrandomForestSRCsanitizer issue (a 0-length array access inrfsrcGrow,entry.c:184) that surfaces when anyvarProgrow (varpro(),beta.varpro(),uvarpro(),isopro(),ivarpro()) builds a forest. ggRandomForests is pure R and its code is unchanged; the varPro tests still run in our CI (the workflows setNOT_CRAN=true) and locally; they are skipped only on CRAN’s check machines, including the gcc-UBSAN check. The upstream issue has been reported to the randomForestSRC maintainers. - The
varprovignette now loads every varPro fit from a precomputed file (vignettes/varpro_precomputed.rds, built byvignettes/precompute_varpro.R), so the vignette performs no live varPro grow duringR CMD check. This removes the same upstream sanitizer path from the vignette build and trims check time. Each chunk falls back to a live fit if the precomputed object is absent, so the vignette remains reproducible from source.
ggRandomForests v3.1.0
CRAN release: 2026-06-11
- Fix:
gg_vimp()for single-outcome rfsrc forests now correctly flags variables with non-positive VIMP in thepositivecolumn (affecting plot coloring). The column was namedVIMP(uppercase) in single-outcome fits but the flag check accessed$vimp(lowercase), leavingpositivestuck atTRUEfor all variables. Surfaced by the Copilot review on PR #109. - Documentation pass. Deepened the varPro-family and rfsrc importance/partial/survival help pages against the upstream randomForestSRC and varPro documentation, and made the line between
gg_vimp()(permutation, Breiman-Cutler importance) andgg_varpro()(varPro release-rule importance) explicit and cross-linked. Vignette prose deepened with the same framing; one-line code-comment fixes; fixed a stale@returningg_roc()(documented ayvarcolumn the function does not return). No user-facing behavior change. - Vignettes: the regression and survival partial-dependence surfaces are now rendered as static
ggplot2heat maps instead of interactiveplotlywidgets, and figures render at 96 dpi. This cuts the installed size from ~17 MB to ~5 MB (theplotlylibrary is no longer bundled into the vignette HTML).plotlyis dropped fromSuggests. - Check time: reduced the
R CMD checkvignette-rebuild and test timings to bring the overall CRAN check comfortably under budget (CRAN flagged the overall check time on the 3.1.0 submission). The regression and survival vignettes use lighter forests (ntree200 / 150, imputationntree100) and coarser partial-dependence grids. The varpro vignette’s threegg_partial_varpro()calls and the Bostonbeta.varpro()fit (~34 s combined) are precomputed offline byvignettes/precompute_varpro.Rand loaded fromvignettes/varpro_precomputed.rds, with an automatic live-computation fallback if the file is absent. Thegg_udependent()tests memoise the per-fit entropy matrix (varPro::get.beta.entropy(), ~1.5 s and a pure function of the fit) instead of recomputing it once per test. No user-facing behavior change.
ggRandomForests v3.0.0
-
Version jump to 3.0.0. The varPro integration is a major scope expansion plus the
gg_partialpro()soft-deprecation, which is major-version territory. Survival / multivariate varPro families, ROC confidence intervals, and hazard estimates are deferred to v3.1.0. - CRAN-audit cleanup: the
gg_brier()/plot.gg_brier()examples move from\dontrunto\donttest(so they execute underR CMD check --as-cranand on CRAN;library(survival)added soSurv()resolves), the per-variablemessage()in the deprecatedsurv_partial.rfsrc()is removed (its one behavior change: that function no longer prints a line per variable), and the README points to the new “varpro” vignette. - Fix: importance plots now consistently put the most-important variable at the top.
gg_varpro(),gg_beta_varpro(), andgg_ivarpro()previously built theirvariablefactor with descending levels, so aftercoord_flip()the most-important variable landed at the bottom — inverted relative togg_vimp(). All three now reverse the factor levels to match thegg_vimpconvention (and thevarImpPlot/vipstandard). Row order andsummary()output are unchanged (still most-important first). A new cross-function test pins the convention. - New vignette: “Exploring variable importance with varPro.” Walks the full gg_* varPro layer (gg_partial_varpro, gg_varpro, gg_udependent, gg_isopro, gg_beta_varpro, gg_ivarpro) on three worked examples — regression (Boston), classification (iris binary + multi-class), and survival (PBC). Includes a family-support matrix documenting which wrapper works for which forest family. Headline document for v3.0.0.
-
gg_ivarpro()andplot.gg_ivarpro(): tidy wrapper and per-variable-distribution / per-observation-profile plots forvarPro::ivarpro()(individual / local variable importance) across regression and classification (binary + multi-class) families. The long-format tidy frame is(obs, variable, local_imp, selected)for regression; classification adds aclasscolumn. NA cells are filtered out and sparsity is surfaced in provenance.which_obs(integer index) collapses to a single-observation profile; the plot switches from a jittered distribution view to a horizontal bar chart.which_class(response level name) collapses to a single class panel; binary fits default to the last factor level (positive class).cutoffacceptsNULL(per-class mean), a scalar, or a named numeric vector — matching the gg_beta_varpro classification contract. Optionalivarpro_fitargument lets callers cache the expensiveivarpro()call. Last of four Phase 4 sub-projects. -
gg_beta_varpro()adds varPro classification support (binary + multi-class). Binary fits default to a single positive-class panel (last factor level); multi-class fits return a long-format frame with aclasscolumn and plot asfacet_wrap(~ class). Optionalwhich_classselects a single class;cutoffaccepts a scalar or per-class named vector. Variables are stored as a factor whose levels are set bymean(|sum-of-class-beta|)descending so every facet shows rows in the same order. Motivating use case: 30-day mortality. - Provenance shape change for
gg_beta_varpro():attr(*, "provenance")$cutoffis now always a named numeric vector — length 1 named"regr"for regression, length K named with the response factor levels for classification. Downstream tooling should read it as a vector and select by name; the prior scalar shape is gone. -
gg_beta_varpro()andplot.gg_beta_varpro(): tidy wrapper and default horizontal bar chart forvarPro::beta.varpro()— the per-rule lasso-β refinement of variable importance. Aggregates per-rule β̂ by variable intobeta_mean = mean(|β̂|)and flags variables above a selection cutoff (defaultmean(beta_mean)). Optionalbeta_fitargument lets callers compute the expensivebeta.varpro()step once and reuse the result across multiple wrapper calls (different cutoffs, snapshot rebuilds, vignette knits).print/summary/autoplotS3 companions follow the existinggg_*conventions. Regression family only — classification, regr+, and survival are tracked under Phase 4d (see the spec for the endpoint map). Third of three Phase 4 sub-projects. -
gg_isopro()gains anewdataargument so a fittedvarPro::isopromodel can score new observations into the same tidygg_isoproframe. Internally the wrapper callspredict.isopro()twice: withquantiles = FALSEto populate thecase.depthcolumn (varPro’s native polarity, lower = more anomalous) and withquantiles = TRUEto computehowbad = 1 - quantile(the wrapper convention, higher = more anomalous). Both polarities are visible in the returned data frame, and the relationship is named in the roxygen. Theplot/print/summary/autoplotS3 companions work unchanged on the new tidy frame; to overlay training and test scores, bind the two extractor calls with amethodlabel column and pass the result toplot(). Second of three Phase 4 sub-projects. -
Fix (gg_isopro training-path polarity). Bug in the original
gg_isopro(PR #94): varPro’s$howbadon anisoprofit uses “lower = more anomalous” polarity (it is the quantile ofcase.depth), but the wrapper’s plot method and documentation both assume “higher = more anomalous”. Train scores and the new test-data scores were anti-correlated until this PR’s training-path flip (howbad = 1 - object$howbad) brought them into agreement. The fix surfaced because the test-data sanity check (training-as-newdata top-5 overlap) failed at 0/5 instead of 5/5 before the flip. Note: the two vdiffr baselines recorded in PR #94 (gg-isopro-defaultandgg-isopro-threshold) were recorded under the inverted polarity; they are visually flipped relative to the new behavior but CI skips snapshots (VDIFFR_RUN_TESTS = false) so no failure surfaces. Re-record withVDIFFR_RUN_TESTS = truewhen convenient. - Documentation: pedagogical pass over the varPro wrappers (
gg_partial_varpro,gg_varpro,gg_udependentand theirplot.*methods). Each help page now has explicit “What X is doing”, “What’s in the output”, and “What you use this for” sections so a reader new to varPro can learn the underlying method (release rules, beta-entropy dependency, parametric / nonparametric / causal partial estimators) from the help page alone, not just the wrapper mechanics. No API or behavioral change. - Documentation: enable roxygen2 markdown package-wide via
Roxygen: list(markdown = TRUE)inDESCRIPTION. New roxygen blocks can use backticks and[fn()]link syntax; existing\code{}/\link{}markup keeps working. Two source-roxygen edits to keep R CMD check clean:randomForest[SRC]inR/help.R(markdown read it as an unfinished link) becomes plainrandomForestSRC; the95\%escape inR/gg_rfsrc.R::bootstrap_survivalbecomes a literal95%. No API or rendered-doc behavioral change beyond the conventions switch. - New
gg_isopro()andplot.gg_isopro(): tidy wrapper and ranked-elbow + density visualization forvarPro::isoproisolation-forest anomaly scores.plot.gg_isopro()takespanel = c("both", "elbow", "density")and optionalthreshold(score-space) ortop_n_pct(quantile-space) to draw a reference line; if both are set,thresholdwins with a message. Amethodcolumn auto-triggers color grouping for multi-method comparisons (usedplyr::bind_rows()on threegg_isopro()calls).print/summary/autoplotS3 companions follow the existinggg_*conventions. First of three Phase 4 sub-projects. -
plot.gg_variable(): fix render error on the default multi-class classification plot. The default-xvar selection was treatingyvar(the observed-class column) andoutcome(the multi-class pivot facet) as predictors; pivoting them intovarthen dropped the column the downstreamgeom_jitter(aes(color = yvar))referenced, and the patchwork errored when actually rendered. CI did not catch this because the existing test only asserted the patchwork class (lazy) and snapshots run withVDIFFR_RUN_TESTS = false. New test exercises a real build of every sub-plot. -
plot.gg_variable(): the same default-xvar selection used substringgrep("time", ...)/grep("event", ...), which silently dropped any predictor whose name contained those substrings – e.g. the documented veteran-data survival predictordiagtime. Switch to exact matching forevent/time/yvar/outcomeand an anchored prefix foryhat(yhatoryhat.<class>). New test exercisesdiagtimeon the veteran survival forest. -
gg_roc(): per-class one-vs-rest ROC curves (#88, closes #72).- New
per_classargument, defaultFALSE. Withper_class = TRUEon a forest of more than two classes,gg_roc()returns a long-formatgg_rocdata frame with aclassfactor column, plus a named AUC vector attribute with one entry per class, ordered by descending AUC. -
plot.gg_roc()gainspanel = c("overlay", "facet"). When the object has aclasscolumn,"overlay"colors the curves by class and"facet"gives each class its own panel. -
summary.gg_roc()prints the named per-class AUC values when aclasscolumn is present. - On a binary forest,
per_class = TRUEdoes nothing, the usual single-curve result comes back unchanged. - ROC confidence intervals are still to come, in v3.1.0 (issue #7 / #72-CIs).
- New
- New
gg_udependent(): varPro cross-variable dependency (Phase 3).-
gg_udependent()reads cross-variable dependency scores off auvarprofit, viavarPro::get.beta.entropy()andvarPro::sdependent(). It returns a tidy list:$edges(variable_from, variable_to, weight),$nodes(variable, degree, selected), and$graph, an igraph object. -
plot.gg_udependent()draws the dependency network with ggraph. Edge width and opacity scale with dependency strength; node color marks the signal variables. The layout is configurable ("fr","kk","stress", and so on). -
ggraphadded toSuggests:.
-
- New
gg_varpro(): varPro variable importance (#85).-
gg_varpro()pulls per-tree importance scores from a fittedvarproobject and draws a boxplot of the per-tree z-score distribution for each variable. The hinges sit at the 15th and 85th percentiles and the whiskers at the 5th and 95th, so the box is not the usual Tukey one — it reports the percentiles it actually shows. Variables with aggregate z abovecutoff(default 0.79) are color-highlighted. - With
faithful = TRUE, the individual per-tree z-scores are jittered over the box as semi-transparent points, with a white-outlined dot at the mean, the same view as varPro’s internalbxpoutput. - With
conditional = TRUE(classification forests only),gg_varpro()reads$conditional.zand draws class-conditional importance as afacet_wrap(~class, nrow=1)bar chart. - Set
local.std = FALSEto allowplot(..., type = "raw"), which shows raw per-tree importance instead of the z-normalized values.
-
-
gg_variable.randomForest: classification fix (#87).- For a classification forest,
gg_variable.randomForest()now stores per-class OOB vote fractions asyhat.<classname>columns, read fromobject$votes, the same layout therfsrcpath produces. It used to store a singleyhatfactor column of class labels (fromobject$predicted), and that column shape stopped the multi-class pivot inplot.gg_variablefrom ever running. The vote fractions are row-normalized to[0, 1], even when the forest was fit withnorm.votes = FALSE. -
plot.gg_variable, binary classification: withsmooth = TRUEthe x and y aesthetics are now mapped onto the smooth layer correctly. -
plot.gg_variable, multi-class numeric path:smooth = TRUEnow adds the smooth layer instead of skipping it silently. - Closes stale issues #81 (fixed in PR #83) and #82.
- For a classification forest,
- New
gg_partial_varpro(): varPro partial dependence (#84).-
gg_partial_varpro()takes over fromgg_partialpro()as the entry point for varPro partial dependence plots. It accepts an optionalobjectargument (the originatingvarprofit) which it uses for provenance-aware axis labels, and ascaleargument ("auto","mortality","rmst","surv","chf"). - Ensemble mortality labeling (Ishwaran et al. 2008): with
scale = "mortality", orscale = "auto"on a survival forest, the y-axis reads “Ensemble mortality (expected events)”. That is an unbounded relative-risk score, not a survival probability, and the documentation says so plainly so it is not misread. - Survival path C: with
scale = "surv"orscale = "chf",gg_partial_varpro()pulls the embedded rfsrc forest fromobject$rfand returns true S(t) or CHF partial curves through the existinggg_partial_rfsrcmachinery. -
varProis now a hard dependency (Imports:). -
gg_partialpro()is soft-deprecated: it warns, then hands off togg_partial_varpro(). It will be removed in the release after v3.0.0.
-
- randomForest engine validation and repair (#82). Fixes #80, #81, and a
plot.gg_errorlabel wart, and adds full randomForest regression test coverage. Details below.-
plot.gg_variable()now always returns a singleggplot(one variable) or apatchworkcomposite (several variables, or the default) — never a bare list. This matches the v2.7.3plot.gg_partial*change. A list used to come back for multiplexvar, which brokepatchwork/autoplot()/layer_data()composition (#80). -
gg_roc()andcalc_roc()forrandomForestnow build the ROC from class probabilities (OOB votes by default, honoringoob) rather than the degenerate three-point curve they produced before. Withwhich_outcome = "all"(the default forgg_roc(rf)) the result is a macro-averaged one-vs-rest ROC, and no warning. The shared.validate_which_outcomehelper andcalc_roc.rfsrcare byte-for-byte unchanged, so rfsrc behavior is untouched (#81).
-
- Dependency modernization. This breaks scripts that relied on attachment.
randomForestSRCandrandomForestmove fromDepends:toImports:;igraph,callr, andvarProare added toSuggests:(varProlater moves up toImports:, with the first varPro-integration component).library(ggRandomForests)no longer putsrandomForestSRCorrandomForeston the search path. A script that calledrfsrc()orrandomForest()unqualified after onlylibrary(ggRandomForests)now needs its ownlibrary(randomForestSRC)/library(randomForest), or must qualify the calls. ggRandomForests itself is unaffected. It qualifies every call into its dependencies.
ggRandomForests v2.7.3
CRAN release: 2026-05-12
-
plot.gg_partial(),plot.gg_partial_rfsrc(), andplot.gg_partialpro()now always return a singleggplot/patchworkobject. Previously, when both continuous and categorical predictors were present, they returned a named listlist(continuous=, categorical=), which surprised users and madeautoplot()dispatch ambiguous. The two panels are now combined vertically viapatchwork::wrap_plots()(patchwork moved fromSuggeststoImports). Closes #77. -
autoplot()S3 methods for all 10gg_*classes, delegating to the correspondingplot.gg_*()method so objects work in|>pipelines,patchwork, andcowplotcompositions viaggplot2::autoplot(). -
print()andsummary()S3 methods for everygg_*data object (gg_error, gg_vimp, gg_rfsrc, gg_variable, gg_partial, gg_partial_rfsrc, gg_partialpro, gg_roc, gg_survival, gg_brier).print()is header-only — usehead()for rows.summary()returns a printablesummary.ggobject with per-class diagnostics. Eachgg_*constructor now attaches a"provenance"attribute (source, family, ntree, n, xvar.names) consumed by the new methods. - New
gg_brier()extractor andplot.gg_brier()method for time-resolved Brier scores and CRPS on survival forests (issue #9). WrapsrandomForestSRC::get.brier.survival()and adds the mortality-quartile decomposition, a 15-85 percent per-subject envelope, and running CRPS via trapezoidal integration. Supportscens.model = c("km", "rfsrc"),type = c("brier", "crps"), andenvelope(overall line + 15-85% ribbon). Multi-model comparison is left todplyr::bind_rows()on multiplegg_brieroutputs — see?gg_brierfor an example. - Visual unification of ribbon overlays across plot methods. All ribbons now use a shared alpha (
.gg_ribbon_alpha = 0.2) and a shared fill (.gg_ribbon_fill = "steelblue") for single-series cases (KM/NA CIs, bootstrap CIs,gg_brierenvelope); group-stratified ribbons keep their group-colored fill. Statistical bounds unchanged — only styling. ggRandomForests v2.7.2 ===================== - Address CRAN reviewer (Benjamin Altmann) feedback on the v2.7.1 resubmission:
- Add methods references to
DESCRIPTION(Breiman 2001 and Ishwaran et al. 2008, with<doi:...>auto-links) per CRAN cookbook. - Drop the
man/shift.RdRd file:shift()is an internal utility and the example usedggRandomForests:::shift(...). Marked the function@noRdso it no longer generates a help page. - Replace
cat()insurv_partial.rfsrc()withmessage()so progress output is suppressible (suppressMessages()) and plays nicely inside notebooks / Shiny / quarto. - Restore the user’s
par()settings in thesurv_partial.rfsrc()example viaoldpar <- par(no.readonly = TRUE); on.exit(par(oldpar)).
- Add methods references to
ggRandomForests v2.7.1
- Fix
gg_partial_rfsrc()for survival forests:partial.rfsrc()was being called withoutpartial.type, causing a zero-length comparison (if (partial.type == "rel.freq") ...) inside the C-level prediction routine and aborting the call. Survival forests now passpartial.type = "surv"(default; configurable via the newpartial.typeargument accepting"surv","chf", or"mort"). This unblocks thepartial-depchunk in the survival vignette. - Fix
gg_partial_rfsrc()for survival forests with multiplepartial.timevalues:get.partial.plot.data()returns yhat as an[length(partial.values) x length(partial.time)]matrix, but the previous code assumed a vector and crashed on column-mismatch when assigningtime. The result is now reshaped to long form so each(x, time)pair is a single row. - Improve
plot.gg_partial_rfsrc()survival layout: predictor value is now on the x-axis with one curve per (rounded) time point colored byTime, faceted by variable name. The previous default put time on the x-axis and one curve per predictor value, producing a saturated legend with dozens of nearly-identical lines. - Add
tests/testthat/test_plot_layer_data.R: regression suite that usesggplot2::layer_data()to verify eachplot.gg_*()method renders non-empty layers for every supported forest family. Catches the empty-figure class of bug (transform/plot column-name mismatch) without requiring visual inspection. -
ggrandomforests.news()now readsNEWS.md(the canonical change log R also surfaces viautils::news()). The legacy hand-maintainedinst/NEWShas been removed — it had silently drifted to v2.4.0 (June 2025) across three releases, so users running the helper saw stale version info. One source of truth, no more drift window. - Fix
plot.gg_vimp()legend duplication: the bar geom mapped bothfillandcolorto thepositivecolumn, but only the fill legend was titled “VIMP > 0”, leaving a redundant second legend titled “positive”. Both aesthetics now share the “VIMP > 0” title so ggplot merges them into a single legend by default. - Fix
plot.gg_vimp()for forests with all-positive VIMP: the bar geom previously mapped onlycolor(nofill), producing hollow / outline- only bars and an “Ignoring unknown labels: fill” warning wheneverlabs(fill = ...)was applied. Bothfillandcolorare now mapped unconditionally, so bars render filled in every case. - Add
@examplesblocks toplot.gg_partial_rfsrc()andplot.gg_partialpro(). The latter uses a self-contained mock of thevarpro::partialpro()output structure so the example runs without pulling invarproas a dependency.
ggRandomForests v2.7.0
- S3 design overhaul:
gg_partial(),gg_partialpro(), andgg_partial_rfsrc()now stamp their return values with S3 classes (gg_partial,gg_partialpro,gg_partial_rfsrcrespectively), enablingplot()dispatch without any boilerplate. - Add
plot.gg_partial(),plot.gg_partial_rfsrc(), andplot.gg_partialpro()S3 methods; continuous predictors render as line plots, categorical as bar charts, faceted by variable name. Survival forests produce curves over time; two-variable surface plots group byxvar2.name. - Convert
gg_survival()to an S3 generic dispatching on the class of its first argument. Newgg_survival.rfsrc()method extracts the survival response directly from the fitted forest (no separate data argument needed);gg_survival.default()preserves the existing interface. - Fix
plot.gg_survival()auto-coercion: previously calledgg_survival(rfsrc_obj)treating the forest as theintervalstring argument, causing a latent crash; replaced withinherits()guard. - Deprecate
surv_partial.rfsrc()via.Deprecated()with a pointer togg_partial_rfsrc(); all package tests updated to suppress the warning. - Fix
gg_partial_rfsrc()—make_eval_grid()usedunlist(dplyr::select())which coerced factor columns to integer codes; now usesnewx[[xname]]to preserve column class. Categorical detection extended to coveris.factor()andis.character()in addition to the cardinality check. - Add guards to
gg_partial_rfsrc(): all-NAxvalafter NA removal now emits a warning and skips the variable; all-NA grouping variable (xvar2) callsstop();n_evalandcat_limitare validated as single integers >= 2 near function entry. - Fix cyclomatic complexity across
gg_partial_rfsrc.R: refactored into eight top-level unexported helpers (validate_scalar_int,validate_partial_args,snap_partial_time,make_eval_grid,call_partial_rfsrc,partial_one_var,partial_no_group,partial_with_group,split_partial_result); all functions now score below thecyclocomp_linterlimit of 20. - Fix
@param partial.timedocumentation: “see the section above” corrected to “see the section below”. - Replace deprecated
tidyr::gather()withtidyr::pivot_longer()inplot.gg_vimp()andplot.gg_partialpro(). - Add
gg_survival.rfsrc,gg_survival.default,plot.gg_partial,plot.gg_partial_rfsrc, andplot.gg_partialprotoNAMESPACE; add corresponding@rdname/@exportroxygen tags. - Update tests: add
expect_s3_class()checks for all new classes; addplot()smoke tests forgg_partial,gg_partial_rfsrc,gg_partialpro; addgg_survival.rfsrctests for KM extraction,bystratification, and error on non-survival forest. - Add
plot.gg_partial,plot.gg_partial_rfsrc, andplot.gg_partialproto_pkgdown.ymlreference index.
ggRandomForests v2.7.0
- Fix critical visual bug in
plot.gg_rfsrc: allaes()calls used bare string literals instead of.data[[col]], causing every aesthetic to map to a constant string rather than the underlying data column. All plot types (regression, classification, survival) were affected. - Fix
aes()bare-string literals inplot.gg_rocmulti-class branch; remove unreachableif (crv < 2)dead-code branch. - Fix
bootstrap_survivalCI-band indexing ingg_rfsrc: negative index computed viacolnames()was a no-op on large datasets and a latent crash for data with ≤ 2 unique event times. - Fix
gg_rfsrc.rfsrc:is.null(df[, col])does not detect missing columns; replaced with!col %in% colnames()guard. - Fix
gg_rfsrc.randomForest: method used non-existentobject$xvar; now recovers the training frame via.rf_recover_model_frame(). - Fix legend suppression in
plot.gg_errorfor single-outcome forests where the data frame has novariablecolumn. - Fix
gg_vimpandplot.gg_vimp:1:nvarreplaced withseq_len(nvar)in both S3 methods;1:0silently returnedc(1, 0)instead ofinteger(0)whennvar == 0. - Migrate full test suite to testthat 3.x API:
expect_is→expect_s3_class/expect_type/expect_true(is.*());expect_equivalent→expect_equal(ignore_attr = TRUE); allcontext()calls removed; testthat 1.xexpect_that/is_identical_toremoved. - Add
.lintrpackage-level linter configuration; fix lintr spacing ingg_partial. - Improve GitHub Actions:
lint.yamlnow fails CI on any lint issue;R-CMD-check.yamltreats warnings as errors and uses Rtools 44;test-coverage.yamlduplicate codecov upload removed. - Add
covrandvdiffrtoSuggests.
ggRandomForests v2.6.1
- Fix model-label assignment in
gg_partialfor categorical variable data - Refactor
gg_partialandgg_partial_rfsrcto improve factor-level normalization and categorical data handling
ggRandomForests v2.6.0
- Add and export new plotting functions; update existing plot documentation
- Improve unit and integration tests; overall coverage raised to 83%
- Remove
hvtiRutilitiesinternal dependency; clean up associated imports - Refactor
gg_partial_rfsrcto use.datapronoun for alldplyrcalls
ggRandomForests v2.5.0
- Initial
gg_partial_rfsrcfunction: computes partial dependence data directly from anrfsrcmodel viarandomForestSRC::partial.rfsrc, without requiring a separateplot.variablecall - Add support for a grouping variable (
xvar2.name) ingg_partial_rfsrc - Improved vignette formatting and namespace usage
ggRandomForests v2.4.0
- Updating to latest ggplot2 functions
- Utilize some namespace referencing
- Added pkgdown documentation
- Minor testing improvements
ggRandomForests v2.2.0
CRAN release: 2022-05-09
- Bring back the regression vignette
- Improve package tests and code coverage
- Clean up code with lintr
ggRandomForests v2.1.0
CRAN release: 2022-04-26
To pull this out of archive on randomForestSRC 3.1 build release. Fixed a plot bug for gg_error to show the actual curve (issue 35)
ggRandomForests v2.0.1
CRAN release: 2016-09-07
- Correct a bug in survival plots when predicting on future data without a known outcome.
- All Vignettes are now at https://github.com/ehrlinger/ggRFVignette
- All tests are being moved to https://github.com/ehrlinger/ggRFVignette
- Begin work on rewriting all checks to not use cached data. This will require more runtime, and hence we will run fewer of them on CRAN release.
- Minor bug and documentation fixes.
ggRandomForests v2.0.0
CRAN release: 2016-06-11
- Added initial support for the randomForest package
- Updated cache files for randomForestSRC 2.2.0 release.
- Remove regression vignettes to meet CRAN size limits. These remain available at the package source https://github.com/ehrlinger/ggRandomForests
- Minor bug and documentation fixes.
ggRandomForests v1.2.1
CRAN release: 2015-12-12
- Update cached datasets for randomForestSRC 2.0.0 release.
- Correct some vignette formatting errors (thanks Joe Smith)
ggRandomForests v1.2.0
CRAN release: 2015-11-15
- Convert to semantic versioning http://semver.org/
- Updates for release of ggplot2 2.0.0
- Change from reshape2::melt dependence to tidyr::gather
- Optimize tests for CRAN to optimize R CMD CHECK times.
ggRandomForests v1.1.4
CRAN release: 2015-03-29
combine.gg_partialbug when giving a single variable plot.variable object.Remove
dplyrdepends to transitions from “Imports” to “Suggests”.Argument for single outcome
gg_vimpplot for classification forests.Improvements to
gg_vimparguments for consistency.Add bootstrap confidence intervals to
gg_rfsrcfunction.Initial
partial.rfsrcfunction to replace therandomForestSRC::plot.variablefunction.Move cache data to
randomForestSRCv1.6.1 to take advantage ofrfsrcversion checking between function calls.Vignette updates for JSS submission of “ggRandomForests: Exploring Random Forest Survival”.
Vignette updates for arXiv submission of ggRandomForests: Random Forests for Regression
Some optimizations to reduce package size.
Remove all tests from CRAN build to optimize R CMD CHECK times.
Remove pdf vignette figure from CRAN build.
Return S3method calls to NAMESPACE for “S3 methods exported but not registered” for R V3.2+.
Misc Bug Fixes.
ggRandomForests v1.1.3
CRAN release: 2015-01-08
- Update “ggRandomForests: Visually Exploring a Random Forest for Regression” vignette.
- Further development of draft package vignette “Survival with Random Forests”.
- Rename vignettes to align with randomForestSRC package usage.
- Add more tests and example functions.
- Refactor
gg_functions into S3 methods to allow future implementation for other random forest packages. - Improved help files.
- Updated DESCRIPTION file to remove redundant parts.
- Misc Bug Fixes.
ggRandomForests v1.1.2
CRAN release: 2014-12-25
- Add package vignette “ggRandomForests: Visually Exploring a Random Forest for Regression”
- Add gg_partial_coplot, quantile_cuts and surface_matrix functions
- export the calc_roc and calc_auc functions.
- replace tidyr function dependency with reshape2 (melt instead of gather) due to lazy eval issues.
- reduce dplyr dependencies (remove select and %>% usage for base equivalents, I still use tbl_df for printing)
- Further development of package vignette “Survival with Random Forests”
- Refactor cached example datasets for better documentation, estimates and examples.
- Improved help files.
- Updated DESCRIPTION file to remove redundant parts.
- Misc Bug Fixes.
ggRandomForests v1.1.1
CRAN release: 2014-12-13
Maintenance release, mostly to fix gg_survival and gg_partial plots. * Fix the gg_survival functions to plot kaplan-meier estimates. * Fix the gg_partial functions for categorical variables. * Add some more S3 print functions. * Try to make gg_functions more consistent. * Further development of package vignette “Survival with Random Forests” * Modify the example cached datasets for better estimates and examples. * Improve help files. * Misc Bug Fixes.