Skip to content
Open
Show file tree
Hide file tree
Changes from 1 commit
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -20,7 +20,9 @@
package org.grails.datastore.gorm.transform;

import java.util.ArrayList;
import java.util.Collections;
import java.util.HashMap;
import java.util.IdentityHashMap;
import java.util.List;
import java.util.Map;

Expand All @@ -47,7 +49,28 @@
* @since 6.1
*/
public class AstPropertyResolveUtils {
protected static Map<String, Map<String, ClassNode>> cachedClassProperties = new HashMap<>();

/**
* Cache of resolved properties per {@link ClassNode}.
* <p>
* Keyed by {@code ClassNode} identity rather than name. {@link ClassNode#equals(Object)} and
* {@link ClassNode#hashCode()} compare by {@link ClassNode#getText()} (essentially the class
* name), so a {@code Map} keyed by name - or even by {@code ClassNode} itself as the map key -
* treats any two distinct {@code ClassNode} instances that happen to share a name as the same
* cache entry. That collision is a real hazard for classes compiled without a package (common
* in tests and dynamically generated sources), and for the same source compiled more than once
* in separate {@code GroovyClassLoader}s: each compilation produces its own {@code ClassNode}
* instance that must never share cached property data with another compilation's instance of a
* same-named class. An {@link IdentityHashMap} avoids that collision entirely by comparing keys
* with {@code ==} instead of {@code equals()}.
* <p>
* Wrapped in {@link Collections#synchronizedMap(Map)} because AST transforms that populate and
* read this cache can run concurrently on multiple threads (e.g. parallel test execution within
* one JVM/fork); a plain, unsynchronized {@link HashMap} is not safe for concurrent structural
* modification and can corrupt its internal state under concurrent {@code put()} calls.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This rationale is inaccurate in a way that will mislead the next reader. maxParallelForks = configuredTestParallel (gradle/test-config.gradle:86) forks separate JVMs, and each JVM gets its own copy of a static field — parallel test forks can therefore never race on this map. JUnit's in-JVM parallel execution isn't enabled anywhere in the build either.

The concurrency exposure that does exist is the compiler itself: multiple compileGroovy tasks running concurrently in a shared Gradle worker, and any embedded compilation driven from more than one thread. Worth rewording to that, otherwise the comment justifies the synchronization with a scenario that can't happen.

@borinquenkid borinquenkid Jul 31, 2026

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You're right, fixed (078ac9d). The rationale no longer claims parallel test forks race on this - confirmed against gradle/test-config.gradle that maxParallelForks really does fork separate JVMs and this build never enables JUnit's in-JVM parallel execution, so that was never a real race. Rewrote the javadoc around the actual exposure instead: interned ClassHelper singletons (OBJECT_TYPE, STRING_TYPE, etc.) that any concurrently-running compilation in the same JVM could resolve to and touch through this cache - which is also why access is now synchronized per-node (c3a4d92) rather than resting on "only one thread ever touches a given node," which turned out not to be universally true.

*/
protected static final Map<ClassNode, Map<String, ClassNode>> cachedClassProperties =

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two breaking changes to a protected member of a public class on one line: the key type changed (String -> ClassNode) and the field became final. Anything downstream that reassigned it — today the only way to clear this cache, which plugin AST transforms and test harnesses plausibly do — now fails to compile rather than merely behaving differently.

If it's being broken anyway, take it the rest of the way: make it private static final and expose an explicit, documented clearCache(). That shrinks the exposed surface and gives the retention problem above an escape hatch, instead of leaving protected visibility on a field nobody can usefully touch any more.

@borinquenkid borinquenkid Jul 31, 2026

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Went further than private + clearCache() - the static map is gone entirely (078ac9d), so there's no field left to expose or protect. That's still technically a breaking removal for anyone who referenced the old protected field directly, just a cleaner one than a silent type change. Given it's an internal implementation-detail field with no getter/documented extension use, I'm inclined to leave it out of the upgrade guide, but flagging it in case there's a reason to cover it there.

Collections.synchronizedMap(new IdentityHashMap<>());

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Keying this static map by ClassNode identity does fix the collision, but it converts a bounded cache into a classloader leak.

Under the old String key the map held at most one entry per distinct class name. With ClassNode keys it holds one entry per instance, and a primary ClassNode transitively pins its entire compilation:

ClassNode.getModule()  ->  ModuleNode.getUnit()  ->  CompileUnit.loader  (GroovyClassLoader)

plus ClassNode.clazz directly for resolved nodes. Since nothing ever removes an entry and the field is now static final, every compilation that runs a GORM AST transform inside a long-lived JVM permanently retains that compilation's classloader and every class it loaded. That is not hypothetical: the Gradle Groovy compiler daemon is reused across compileGroovy tasks and across builds, and dev-mode / GroovyClassLoader-driven recompiles land in the same place. The map values are ClassNodes too, so they pin loaders as well — that part pre-existed, but it was O(distinct class names) and is now unbounded in the number of compilations.

Groovy already stores per-node state exactly this way (ClassNode.getModule() above is itself getNodeMetaData(ModuleNode.class)), so the cleanest fix is to drop the static map and hang the cache off the node:

private static final String PROPERTIES_CACHE_KEY = AstPropertyResolveUtils.class.getName() + ".properties";

private static Map<String, ClassNode> getPropertiesFromCache(ClassNode classNode) {
    return classNode.getNodeMetaData(PROPERTIES_CACHE_KEY, cn -> computeProperties(classNode));
}

getNodeMetaData(Object, Function) is available in Groovy 5, is identity-scoped by construction (so the collision this PR fixes cannot arise at all), needs no global lock, and is collected together with the node. Worth deciding explicitly whether the holder should be classNode or classNode.redirect()getModule() uses redirect(), and redirected nodes are the one case where the two differ.

If a process-wide map has to stay for some reason, it needs weak identity keys and/or an explicit eviction point, plus a note documenting the expected lifecycle.

@borinquenkid borinquenkid Jul 31, 2026

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed, and went with your suggested direction: the static map is gone entirely, and the cache is now stored as ClassNode metadata (classNode.redirect().getNodeMetaData(key, fn)), so a cache entry is only reachable through the node it describes and is collected with it. See the updated javadoc on cachedClassProperties's replacement (PROPERTIES_CACHE_KEY) for the full writeup, including why keying on redirect() is safe here (setRedirect() throws for a primary node - i.e. every real caller of this utility, since they're all mid-compilation - so redirect() is just this for the node's whole life in practice) and why access ended up needing explicit per-node synchronization (some real callers can resolve to interned singletons like ClassHelper.OBJECT_TYPE for a plain Object/def-typed property, and those are shared across every compilation in the JVM, not scoped to one thread the way a normal node is). Landed in 078ac9d, tightened further in c3a4d92 after an adversarial self-review turned up the synchronization gap.


/**
* Resolves the type of of the given property
Expand Down Expand Up @@ -94,22 +117,25 @@ public static List<String> getPropertyNames(ClassNode classNode) {
}

private static Map<String, ClassNode> getPropertiesFromCache(ClassNode classNode) {
String className = classNode.getName();
Map<String, ClassNode> cachedProperties = cachedClassProperties.get(className);
Map<String, ClassNode> cachedProperties = cachedClassProperties.get(classNode);
if (cachedProperties == null) {
cachedProperties = new HashMap<>();
Map<String, ClassNode> newProperties = new HashMap<>();
boolean isDomainClass = AstUtils.isDomainClass(classNode);
if (isDomainClass) {
cachedProperties.put(GormProperties.IDENTITY, new ClassNode(Long.class));
cachedProperties.put(GormProperties.VERSION, new ClassNode(Long.class));
newProperties.put(GormProperties.IDENTITY, new ClassNode(Long.class));
newProperties.put(GormProperties.VERSION, new ClassNode(Long.class));
}
cachedClassProperties.put(className, cachedProperties);
ClassNode currentNode = classNode;
while (currentNode != null && !currentNode.equals(ClassHelper.OBJECT_TYPE)) {
populatePropertiesForClassNode(currentNode, cachedProperties, isDomainClass, !isDomainClass);
populatePropertiesForClassNode(currentNode, newProperties, isDomainClass, !isDomainClass);
currentNode = currentNode.getSuperClass();
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Identity keying makes the cache immune to name collisions, but I don't think it removes the nondeterminism that a flake like #16030 needs. populatePropertiesForClassNode consults ClassPropertyFetcher only when classNode.isResolved() (line 172), so an entry is a snapshot of whatever resolution state the node happened to be in at the first lookup, and it is never refreshed afterwards. If the first lookup for a node lands at a different compilation phase between runs, the cached property set still differs between runs — identity keys or not.

That also weakens the root-cause story in the description: compiling the same source twice produces two ClassNodes with identical property sets, so a name-keyed collision between just those two is harmless and can't produce divergent bytecode on its own. The collisions that would actually diverge are with a differently-shaped same-named class, or with a same-named node whose entry was cached at a different resolution state.

Can you pin down which one you observed — e.g. the two colliding class names, or the flake reproducing in a loop pre-fix and not post-fix? The change is an improvement regardless, but if the resolution-state snapshot is the real driver then #16030 comes back and this gets recorded as already fixed.

@borinquenkid borinquenkid Jul 31, 2026

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Wanted to actually test this rather than argue about it. I reverted to the pre-fix cache and ran the exact "same source compiled twice" scenario WhereQueryClosureCaptureSpec's second test exercises, 40 times sequentially in one JVM (simulating many test classes sharing one fork's static state) - it never diverged, even with the old buggy cache. So I can't point to a repro that pins the ~1% CI flake specifically to the name/identity-collision mechanism, and you may well be right that it isn't the actual trigger. I'm keeping the identity-safety fix regardless, since it's an independently real, provable bug on its own terms (the spec's second test shows two distinct, differently-shaped ClassNodes that happen to share a name getting their properties conflated under the old cache) - but I want to be upfront that I haven't confirmed it's the flake's cause, only that it's a real bug. I'll be watching the flaky-test dashboard (#16030) after this merges to see whether WhereQueryClosureCaptureSpec actually clears.

On the isResolved() snapshot mechanism specifically, I checked the Groovy 5.0.7 source directly: ClassNode.clazz has no setter anywhere outside the ClassNode(Class) constructor, and setRedirect() throws a GroovyBugError if called on a primary node - which is what every real caller of this utility passes in, since they're all resolving a class mid-compilation. So for the nodes this cache actually serves, redirect() is simply the node itself for its entire life, and isResolved() can't flip after the node is first cached. I don't think the staleness you're describing can occur for this code path as it's actually used, though I agree the "cache once, forever" design would be fragile if it could - documented the reasoning (and the primary-node guarantee it leans on) in the class javadoc.

} return cachedProperties;
// Publish only once fully populated so a concurrent reader can never observe a
// partially-populated entry for this ClassNode.
cachedProperties = newProperties;
cachedClassProperties.put(classNode, cachedProperties);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The reorder-and-publish-last fix is correct, and synchronizedMap does supply the safe publication it depends on: get and put synchronize on the same mutex, so a reader that observes the entry also observes a fully-populated HashMap. Worth stating that explicitly in the inline comment, because "publish only once fully populated" is only sufficient given that happens-before edge — with a bare HashMap the reordering alone wouldn't have been enough.

One remaining wrinkle: the check-then-act across the get on line 120 and the put on line 136 is not atomic, so two concurrent callers can each build a complete map and the later put wins. Benign here (the maps are equivalent and never mutated after publication), but cachedClassProperties.computeIfAbsent(classNode, ...) on the synchronized wrapper would be atomic and shorter — the tradeoff being that it holds the mutex for the whole superclass walk.

@borinquenkid borinquenkid Jul 31, 2026

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

getNodeMetaData(key, fn) uses computeIfAbsent internally (078ac9d), and I additionally wrapped the call in synchronized (cacheHolder) (c3a4d92) after an adversarial self-review flagged that the backing ListHashMap is explicitly documented as not thread-safe, and that some real callers can land on interned singleton nodes (ClassHelper.OBJECT_TYPE etc.) that more than one concurrent compilation could reach - the "only one thread per node" assumption doesn't hold for those. See the javadoc for the precise scope of what the synchronization does and doesn't cover.

}
return cachedProperties;
}

private static void populatePropertiesForClassNode(ClassNode classNode, Map<String, ClassNode> cachedProperties, boolean isDomainClass, boolean allowAbstract) {
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,101 @@
/*
* Licensed to the Apache Software Foundation (ASF) under one
* or more contributor license agreements. See the NOTICE file
* distributed with this work for additional information
* regarding copyright ownership. The ASF licenses this file
* to you under the Apache License, Version 2.0 (the
* "License"); you may not use this file except in compliance
* with the License. You may obtain a copy of the License at
*
* https://www.apache.org/licenses/LICENSE-2.0
*
* Unless required by applicable law or agreed to in writing,
* software distributed under the License is distributed on an
* "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
* KIND, either express or implied. See the License for the
* specific language governing permissions and limitations
* under the License.
*/
package org.grails.datastore.gorm.transform

import java.lang.reflect.Modifier

import org.codehaus.groovy.ast.ClassHelper
import org.codehaus.groovy.ast.ClassNode
import spock.lang.Specification

/**
* {@link AstPropertyResolveUtils} caches resolved property metadata in a static, process-wide
* map keyed by {@link ClassNode}. Two distinct compilations (e.g. the same source parsed in two
* different {@code GroovyClassLoader}s, as happens for dynamically-generated sources and in
* tests) produce distinct {@code ClassNode} instances that can legitimately share the exact same
* name - {@code ClassNode#equals(Object)} compares by name, so a naive name- or equals()-based
* cache key would conflate them, corrupting the resolved properties of one class with those of
* an unrelated class that happens to share its name. This spec proves the cache keys strictly by
* {@code ClassNode} identity, so same-named-but-distinct class nodes never contaminate each
* other's cached property data.
*/
class AstPropertyResolveUtilsSpec extends Specification {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good to see this class get its first coverage. As written, though, it only exercises hand-built non-domain ClassNodes, so the branches that actually matter to the transform never execute: AstUtils.isDomainClass is false in all three features, which means the injected id/version entries, the hasMany/belongsTo/hasOne handling in populatePropertiesForInitialExpression, the isResolved() / ClassPropertyFetcher path, and the superclass walk are all untested. This PR rewrites the population loop feeding every one of those, so per the repo rule that a touched class gets its behaviour verified they should be covered here — an @Entity-annotated ClassNode with a hasMany initial expression and a domain superclass would reach most of it.

Also missing: a test for the concurrency fix the PR claims. Several threads calling getPropertyNames concurrently on the same node and on distinct nodes, asserting every returned list is complete, would exercise both the synchronized map and the publish-after-populate ordering. As it stands, reverting either half of the concurrency change leaves this spec green.

@borinquenkid borinquenkid Jul 31, 2026

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added both. Domain-class coverage now includes identity/version injection and hasMany/belongsTo/hasOne resolved two ways: via AST initial expressions (a hand-built ClassNode with a crafted MapExpression, matching what static hasMany = [...] compiles to) and via reflection on an already-resolved class (compiled through GroovyClassLoader.parseClass then re-wrapped with ClassHelper.make(), since that's the only way to get isResolved() == true for that branch).

Also added two concurrency tests: many threads each resolving their own distinct, identically-named ClassNode (proves the identity-collision-freedom property holds under concurrent load, not just sequentially), and - after an adversarial self-review pointed out the first test can't exercise any real race, since nothing is shared between the threads - a second test where 32 threads resolve the exact same shared ClassNode concurrently, which is what the synchronized fix mentioned above actually protects. Worth being honest about that second test's limits: I checked whether it fails without the synchronized guard, and it didn't, in 8 runs - the cached computation is deterministic and idempotent, so a black-box return-value test can't reliably force the underlying race into an observably wrong result. Said that directly in the test's comment rather than overclaiming what it proves; the synchronization is justified by ListHashMap's own "not thread-safe" documentation, not by this test having caught a live bug.


void "property lookups for two same-named ClassNodes in different packages do not corrupt each other"() {
given: 'two distinct ClassNodes with the same simple name declared in different packages'
ClassNode first = new ClassNode('org.example.one.Widget', Modifier.PUBLIC, ClassHelper.OBJECT_TYPE)
first.addProperty('color', Modifier.PUBLIC, ClassHelper.STRING_TYPE, null, null, null)

ClassNode second = new ClassNode('org.example.two.Widget', Modifier.PUBLIC, ClassHelper.OBJECT_TYPE)
second.addProperty('weight', Modifier.PUBLIC, ClassHelper.Integer_TYPE, null, null, null)
Comment on lines +56 to +62

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch on the mechanics — you're right that this test's two ClassNodes use different fully-qualified names (org.example.one.Widget vs org.example.two.Widget), so it wouldn't have collided under the old name-/equals()-keyed cache either, and doesn't by itself guard the regression.

That guard is the next test below, "property lookups for two distinct ClassNode instances with the exact same unqualified name do not corrupt each other" — both ClassNodes there are named plain Widget with no package, so they do compare equal (first == second, matching hashCode()) exactly as the old cache's key would, and the test asserts they still resolve independently under the new identity-keyed cache. That's the one that fails under the old implementation and passes under this fix.

This first test is intentionally a different, narrower check — that two distinct, differently-named classes never get confused with each other, which is a correctness property worth keeping on its own regardless of the collision bug. Leaving both as-is: this one for general non-contamination across genuinely different classes, the next one for the actual same-name collision regression.


when: 'the first class node is resolved, populating its cache entry'
List<String> firstProperties = AstPropertyResolveUtils.getPropertyNames(first)

then: 'only its own property is resolved'
firstProperties.contains('color')
!firstProperties.contains('weight')

when: 'the second, differently-packaged, same-simple-name class node is resolved'
List<String> secondProperties = AstPropertyResolveUtils.getPropertyNames(second)

then: 'its own property is resolved, not leaked from the first class node'
secondProperties.contains('weight')
!secondProperties.contains('color')

and: 'the first class node cache entry remains unaffected by resolving the second'
List<String> firstPropertiesAfter = AstPropertyResolveUtils.getPropertyNames(first)
firstPropertiesAfter.contains('color')
!firstPropertiesAfter.contains('weight')
}

void "property lookups for two distinct ClassNode instances with the exact same unqualified name do not corrupt each other"() {
given: 'two distinct ClassNode instances - as produced by two separate compilations - sharing an identical unqualified name'
ClassNode first = new ClassNode('Widget', Modifier.PUBLIC, ClassHelper.OBJECT_TYPE)
first.addProperty('color', Modifier.PUBLIC, ClassHelper.STRING_TYPE, null, null, null)

ClassNode second = new ClassNode('Widget', Modifier.PUBLIC, ClassHelper.OBJECT_TYPE)
second.addProperty('weight', Modifier.PUBLIC, ClassHelper.Integer_TYPE, null, null, null)

expect: 'the two ClassNode instances compare equal by name - the exact condition that would collide in a name-keyed or equals()-keyed cache'
first == second
first.hashCode() == second.hashCode()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These two assert Groovy's own ClassNode equality contract rather than anything about AstPropertyResolveUtils. They hold today — 5.0.7's ClassNode.equals compares getText() and hashCode() delegates to getText().hashCode() — but if Groovy ever moves ClassNode to identity equality this spec fails while the production behaviour it guards is still perfectly correct.

!first.is(second) on line 79 is the precondition the test actually needs. Consider keeping that and demoting the other two to a comment explaining why two same-named nodes used to collide.

@borinquenkid borinquenkid Jul 31, 2026

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done - trimmed to just !first.is(second), with a comment explaining why the equals()/hashCode() equality matters (it's the exact collision condition a name- or equals()-keyed cache would hit) rather than asserting on Groovy's own ClassNode equality contract.

!first.is(second)

when: 'both class nodes are resolved'
List<String> firstProperties = AstPropertyResolveUtils.getPropertyNames(first)
List<String> secondProperties = AstPropertyResolveUtils.getPropertyNames(second)

then: 'each keeps its own, independently-resolved properties despite comparing equal'
firstProperties.contains('color')
!firstProperties.contains('weight')
secondProperties.contains('weight')
!secondProperties.contains('color')
}

void "getPropertyType resolves and caches the type of a declared property"() {
given: 'a class node with a declared property'
ClassNode classNode = new ClassNode('org.example.PropertyTypeWidget', Modifier.PUBLIC, ClassHelper.OBJECT_TYPE)
classNode.addProperty('label', Modifier.PUBLIC, ClassHelper.STRING_TYPE, null, null, null)

expect: 'the resolved property type matches the declared type, both on first (cache-populating) and second (cache-hit) lookup'
AstPropertyResolveUtils.getPropertyType(classNode, 'label') == ClassHelper.STRING_TYPE
AstPropertyResolveUtils.getPropertyType(classNode, 'label') == ClassHelper.STRING_TYPE

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The name says "and caches", but neither assertion can distinguish a cache hit from a miss — on a miss getPropertyType falls through to classNode.getProperty(propertyName) and returns the same STRING_TYPE, so both lines pass with caching entirely disabled.

To actually pin the caching behaviour: resolve once, then add a second property to the ClassNode and assert getPropertyNames still returns the stale view. That snapshot-at-first-lookup semantics is what this class really guarantees and what the transform relies on, and it's currently unverified.

@borinquenkid borinquenkid Jul 31, 2026

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Replaced it. The new test adds a second property to the ClassNode after the first getPropertyNames lookup and asserts the second lookup doesn't see it - getPropertyNames has no live-fallback path (unlike getPropertyType, which falls through to a direct classNode.getProperty() lookup when the cache doesn't contain the key), so this actually proves the result was cached rather than recomputed. It also asserts, via a direct classNode.getProperty('extra') != null check, that the property really was added to the underlying node - so the test demonstrates staleness specifically, not just a lookup that happens to return nothing.

}
}
Loading