By Anonymous User
Review Details
Reviewer has chosen to be Anonymous
Overall Impression: Good
Content:
Technical Quality of the paper: Average
Originality of the paper: Yes
Adequacy of the bibliography: Yes, but see detailed comments
Presentation:
Adequacy of the abstract: Yes
Introduction: background and motivation: Good
Organization of the paper: Needs improvement
Level of English: Unsatisfactory
Overall presentation: Average
Detailed Comments:
General Comments
================
I think a taxonomy such as the one being proposed is certainly needed, and the paper has opened my eyes to how other recent works (including my own!) overlook the issue of reasoning and proving capabilities and assurance. I also give kudos for drawing attention to older, more traditional NeSy works, as I believe there is a misconception that neurosymbolic computing is a relatively new field. So in terms of relevance and novelty this survey hits the mark.
However, I'm afraid this paper seems some major revisions before I feel it is ready for publication, but I am enthusiastic in encouraging such revisions. My primary concern is a lack of clarity of the proposed taxonomy, especially with respect to interpreting the ordering and hierarchy. Secondary concerns would be a need to expand on how "eras" are defined, a lack of summarised findings and future work proposals in the concluding chapter, a bit of uncertainty with respect to target audience, and a large collection of typographical issues.
(1) Suitability as introductory text, targeted at researchers, PhD students, or practitioners, to get started on the covered topic.
To me it wasn't clear what level of background knowledge the reader is assumed to have. The authors give a very detailed (and good!) introduction to propositional logic, and skim over neural networks. Nothing wrong with this in itself. Based on this I would assume this paper was written to somebody very familiar with the latter and very unfamiliar with the former, but this is not clarified.
Furthermore, assuming that somebody completely new to propositional logic reads section 2, I get the feeling that some of the terminology introduced in chapter 3 is a bit of a jump. For example terms like "lemma" and "witness" are used in chapter 3 but not defined here or in section 2. Note also my below comment about distinguishing between general and specific cases in the hierarchy - for a novice in propositional logic this may be difficult to appreciate.
(2) How comprehensive and how balanced is the presentation and coverage
A wide range of other reviews are covered, with 43 specific methods are selected in the survey. While not a deal-breaker for me, I couldn't help but feel this felt disproportionately small considering how many methods must be covered by the other reviews overall. Either more papers should be analyses, or a rationale for selection should be clearly specified.
As a positive, papers spanning many years are covered, which is very nice to see considering people sometimes misconceive that neurosymbolic is a new field.
(3) Readability and clarity of the presentation.
Though I nonetheless found the subject matter to very interesting and relevant, I'm afraid to say clarity is my major concern.
Firstly, perhaps some of my confusion would be relieved if I was familiar with the SKOS system, but it is new to me. Authors may need to elaborate on this in their preliminaries if its fundamental to understanding the taxonomy, or clarify that they assume the reader to be familiar with it (see above point about target audience).
In chapter 3 authors discuss the notion of "levels" of capability/assurance. At first when reading chapter 3 I thought "levels" referred to hierarchy depth (e.g. X1.1 is higher level than X1.1.1.) but I didn't appreciate until section 6 that "level" applies to the "a" in "1.a". Section 6, second paragraph: "A higher value of a indicates a higher reasoning capability level". It would have been helpful to have this explained in chapter 3. Anyway, assuming then this means that, for example, X1.8 has higher reasoning capability than X1.1., I'm not really clear on how this aligns with the titles assigned to the various levels. For example, what does it mean for a "robust and interpretable" (X1.8) solution to have higher reasoning capability than a solution that offers "search and control" (X1.4)? This suggests "All Robustness and Interpretability methods enable search and control but not all search and control methods enable robustness and interpretability". Are the authors saying that it is impossible for a model to be robust and interpretable without search and control functionality?
I think it's great that the authors are providing a means to categorise NeSy work according to terms such as robustness, interpretability, search, control, etc., and according to propositional logic more generally, but I'm afraid I just don't understand the ordering of levels.
I'm also a bit unclear on cases where a category of the hierarchy only breaks down into a single sub-group. For example the only child of X1.1.1 Decision Mechanisms is X1.1.1.1."Basic decisions", which implies the existence of "Complex decisions" or "Advanced Decision", which the hierarchy or accompanying text don't address. And again there is X1.1.1.1.1 "Decision-only", suggesting a "Decision-not-only". Maybe 1.1.1.1 and 1.1.1.1.1 are special cases, but the accompanying text doesn't elaborate on the definition of the general case. Any reader new enough to propositional logic to need to read section 2 is going to need help filling in these gaps (also relates to my later comments on target audience).
Some elements of X and Y have counterparts (e.g. X1.7.1 and Y1.1.1). How are they distinct and/or what is the relationship between them? More clear definitions of capability (X) and assurance (Y) and their relationship would help here.
Regarding Z, as I read sections 4 and 5 it wasn't clear to me why Z wasn't included in the analyses of other literature reviews (section 4) and NeSy methods (section 5). I looked at tables 10-16 and wondered where Z was. It only made sense when I started reading chapter 6. Authors should please:
- Clarify in section 3 that Z is not relevant to sections 4 and 5, and why it is not relevant.
- Elaborate on Z more in section 3 (possibly moving all explanation of Z methodology here or at least a summary) - at the very least please explain that it is derived from X and Y.
- Clarify what they mean by "evaluation". In a sense analyses in sections 4 and 5 are also evaluations, just qualitative ones, whereas section 6 is a quantitative evaluation (on data gathered in section 5).
Z1.3: If the process is supposed to be automated (page 22, last paragraph) how does one automatically decide on the choice of {1, 2, 3} for A in equation 13? I'm also not clear how what it means for integration of neural learning and formal inference to be "strong" or "weak". This these points unclear I unfortunately found it difficult to appreciate what any of the discussions on Z1.3 were saying.
I think some improved notation and formatting would also help the taxonomy be more accessible.
- For example, the first digit in any category's enumeration is redundant: X1.a could be X.a and Y1.a could be Y.a with no information loss. With the redundant 1 made the hierarchy feel 1 level deeper than it was, which was a bit intimidating when already trying to navigate a hierarchy of depth 4.
- Breaking down sections 3.1 and 3.2 with further sub-subsections at least for each [X/Y].a would make the explanation of the hierarchy a bit easier to follow. For example a 3.1.1 sub-subsection for "foundational reasoning".
(4) Importance of the covered material to the broader Neurosymbolic AI community.
The paper is very important in this respect. I feel that the issues of validation, proof etc. have been overlooked in recent years, at least with respect to what I've read. There's been disproportionate attention given to explainability and interpretability compared to other properties of neurosymbolic models. This paper was a wake-up call to me at least! This is the main reason I want to re-emphasise that I do strongly encourage the authors to address my concerns, because I would very much like to see this published.
Other general comments:
Eras:
In the background review, I was not clear on the rationale used to partitioned years into eras. For each era could use some justification for the choice of name and perhaps more importantly why one era ends and the next begins. This is important given that they influence the plots in chapter 6, which would appear differently to somebody who may decide to partition the years differently.
That said, the plots in chapter 6 may be better plotted by calculating Z values for each actual year of publication. Given that authors have the years the papers were published they easily have the data available to plot at this level of granularity along the horizontal axis. It will give a more accurate picture of how the Z dimensions change over time that is not impacted by how they or anybody else chooses to partition the different eras and I'd be very interested to see this. They could still plot the eras as annotations over the more fine-grained plot.
Typographical issues:
There are many spelling and grammatical errors throughout. I don't wish to list them all since they may be eliminated once the more major corrections are addressed anyway. But authors should please run a grammar and spell-check before resubmitting.
Section-by-section comments:
=========================
Introduction
Authors should maybe elaborate more on why these guarantees, proving capabilities and assurance levels are important. What are some real-world application areas where this could have critical consequences? This sort of opening discussion helps motivate the reader to keep reading. If people are new to formal systems, what motivates them to learn about it and frame theirs or other NeSy methods in this way?
It's very relevant to recent discussions on guaranteed safe ai, actually (https://arxiv.org/pdf/2405.06624?), where formal guarantees of AI behaviour are encouraged.
Please provide a reference to the SKOS ontology in paragraph 2
Maybe also elaborate more on the scope. Why stop at propositional logics? What about predicate logic? Temporal logic? For example.
Background
Positions of 2.6 and 2.5 should perhaps be swapped. At present 2.5. is introducing neuro-symbolic before neural networks are introduced.
2.1. Propositional Logic: Just want to reemphasise I think this is a good introduction to propositional logic.
2.2. Inference in PL: I'm not sure if both the last paragraph of this section or the first paragraph of 2.3. are necessary as they both say roughly the same thing. I'd lean towards removing the last paragraph of 2.2.
2.6. Short intro to ANNs: Maybe elaborate here that these form the basis of modern AI such as LLMs, CNNs, GNNs, etc.
3. Taxonomy for NeSy Frameworks Survey
Note my above suggestion of using subheadings that reflect the taxonomy, at least at the 1.a level. It might help to make this section easier to digest. It was tricky to get a feel for where I was in the hierarchy as I read through the paragraphs. The redundant 1 after [X/Y/Z] make this a bit more confusing too.
Please introduce the SKOS ontology a bit of an introduction. Why was this chosen as opposed to other taxonomies?
3.1. X: Capability: Page 6, first column, last sentence of penultimate paragraph: please elaborate on global versus local heuristics.
3.2. Y: Assurance
What is an assurance level?
"Interactive explainability (Y1.2.1) is a narrower concept of Explainability for assurance (Y1.2) in the X–Y–Z taxonomy, which generates step-by-step, human-interpretable reasoning traces." - I agree interactive vs non-interactive explanation is an important distinction with significant impact on user experience, but why does explainability have to be interactive to be human-interpretable? Why can't non-interactive explanations be human-interpretable?
Note that by "interactive" I'm imagining a system where a user's own line of enquiry dictates which explanations are displayed, for example the option to expand or collapse an explanation depending on the user's interest. A non-interactive explanation would simply be the output of a report. But both would need to be human-interpretable to be useful. Please clarify what you mean by "interactive" and "human interpretable".
3.3. Z: Evaluating: What is an integration level? What does it mean in practice for one method to be more integrated than another? Can you give examples?
4. Existing Surveys
This section is very interesting and eye-opening but it felt quite long considering the abstract puts focus on the contents of sections 5 and 6. Admittedly I'm not sure what to suggest. Can it be compressed or moved to a later point in the survey?
Tables 2/3: It might be interesting to include the dates of the surveys and to order by date to support later comments about trends with respect to alignment with the dimensions of the taxonomy.
For table 3 ID 18: Ref [123] was not the original proponent of the ADT taxonomy (https://www.sciencedirect.com/science/article/abs/pii/0950705196819204), which dates back as far back as 1995 (during the pioneering era of section 5.2). I'd recommend including the original in the references and some comments on the same in the main text. [123] borrows the taxonomy to frame new ideas from an earlier perspective. Please address this where [123] is discussed in the text also.
5. Literature Survey
Please elaborate on how the 43 papers selected, and why only 43? There's a lot of NeSy methods out there
Please remind the reader why Z is excluded from the tables if section 4 remains where it is at its current length.
As mentioned above please include more commentary on transitions between eras (extends to defending choice of name).
For example, is the message that during 2023-2024 graph-based models were the dominant paradigm for NeSy? Were models before 2018 not differentiable?
5.1. 1943-1989: Foundational neural logic era: While I absolutely agree it's important to address McCulloch and Pitts but I'm not sure we can call this an 'era' when it's just one paper significantly earlier than the pioneering era begins.
5.4. 2018-2020: Differentiable Neuro-Symbolic AI Era: Where does the feature of differentiability come into this? "differentiable" doesn't come up until the last paragraph of the page. Furthermore, older models were differentiable. KBANN and CILP (Pioneering era) were both differentiable. Why do differentiable models stand out during 2028-2020 in particular?
5.6. 2023–2024: Graph-based: A survey on Graph Neural Networks for NeSy was published back in 2020 (differentiable era): https://arxiv.org/pdf/2003.00330. What makes graph-based models stand out during 2023-2024?
Evaluation of Existing Work
I feel it is difficult to include the foundational era in the plots here. With one paper so early on with respect to the busy activity in the other eras, it's really more of an outlier. Don't get me wrong though, it's a very important paper and It's nice to see it discussed earlier.
Note my earlier suggestion to plot against publication date. I appreciate there would probably still be a downward trend, and my intuition from reading so many NeSy papers down the years would lead me to agree with the observation anyway. I think nobody would be able to argue with the trend if it is presented at the finest possible granularity, and the authors have the data to plot it I think.
When discussing the median, what is meant by "central" level of reasoning?
Regarding the decline 2025 for Z1.1, I wonder if th possible effect of citation lag has been considered? Literature that changes the result for 2025 may exist, but because it is new, authors may not be aware of it yet (e.g. because it's too new for other authors to have also discovered and cited it yet). Authors may want to highlight this caveat before drawing conclusions about the numbers for 2025 as a whole.
Conclusions
This section explains what the survey did but unfortunately lacks any commentary on what it concludes. What are the take-home messages? What does the survey show is missing from other surveys and the technologies themselves? How should NeSy researchers take these findings and change how they critique their own and other researcher's work?
I think the authors pick on something quite important in this paper but unfortunately don't give it the emphasis it deserves in the conclusions: that reasoning and proving capabilities and assurance in propositional logic (PL) remains limited in modern NeSy models and surveys, perhaps because of such an emphasis on explainability in the past decade, and more attention needs to be given to these capabilities.
Also is there scope to extend to predicate logics, temporal logics, and other type of logic in general?