Hoojah · Master Research
Online discourse is widely diagnosed as a degradation of the public sphere: flat comment threads, engagement-maximizing ranking, and the collapse of reasoned exchange into reaction. This paper advances the thesis that deliberative quality is a design outcome rather than an emergent accident of a medium, and that a production platform can be built to instantiate deliberative norms directly in its mechanics. We conduct a design analysis of Hoojah, a deployed server-rendered debate platform, reading its mechanisms against two theoretical registers: normative deliberative-democracy theory and formal computational argumentation. On the normative axis we distill five design requirements (reason-giving, structured contestation, equality of standing, considered rather than performed opinion, and early structure against derailment) and show, mechanism by mechanism, how Hoojah's hujah claim tree, three-option stance vote, vote-to-respond gate, phased one-on-one debate, and secret-ballot tally suppression are deliberate instantiations of those requirements. On the formal axis we argue that Hoojah's data model is, by construction, a machine-readable argumentation structure: because a stance vote is a precondition for replying, every argument in the tree carries an explicit, interface-supplied stance label, which converts stance from an inference problem into a recorded datum and renders the graph legible to abstract argumentation semantics. We contribute a mechanism-level mapping from theory to a deployed design, a formalization of the platform's data model as an argumentation framework, and an honest built-versus-proposed ledger paired with a falsifiable evaluation agenda. The paper is forward-looking problem-to-solution research: the literature defines the deficit and the norms; Hoojah's design is offered as a justified solution, with the boundary between what is shipped and what is proposed marked throughout.
The normative promise of networked communication was that it would enlarge the public sphere. In the account that still frames the field, the public sphere is a domain of social life in which private people come together as a public to engage in rational-critical debate, forming opinion through the force of the better argument rather than the weight of status or the reach of a broadcaster [1]. A genuinely open communicative space, on this reading, is one in which claims are advanced and contested on their merits among participants who treat one another as equals. Early enthusiasm held that the internet, by lowering the cost of speaking and widening access, would realize this ideal at scale.
The platforms that actually emerged are, in their dominant design, a structural negation of that ideal. Mainstream social media organize discourse into flat comment threads with no representation of who is arguing with whom or on what side; they rank contributions by predicted engagement rather than by argumentative contribution; and they reward reaction over reason, because the metrics that drive distribution respond to salience and affect. The result is a communicative environment optimized for the circulation of stimulus rather than the exchange of justification. The problem is not that people online refuse to reason. It is that the venues in which they gather are not built to make reasoning visible, structured, or consequential.
Three properties of the dominant design deserve to be named precisely, because each is a lever a different design could pull the other way. The first is the flat thread: a reply sits in an undifferentiated list, with no representation of the stance it takes toward the claim it answers, so that agreement, disagreement, and mere elaboration are visually and structurally indistinguishable. The second is engagement ranking: the order in which contributions are surfaced is governed by predicted attention rather than by argumentative contribution, so that the loudest, most affectively charged, or most socially amplified contribution rises regardless of whether it advances the exchange. The third is the reaction affordance: the cheapest and most rewarded acts are the ones that require no reasons, the one-tap endorsements and the running tallies that convert opinion into a spectacle of social proof. None of these is a law of the medium. Each is a choice, and each admits an alternative choice, which is the premise the rest of this paper builds on.
A substantial body of work has documented the symptoms and, more importantly, established the diagnostic premise on which this paper rests: that the quality of online deliberation is not a fixed property of the medium but a product of design and choice [2]. If deliberative quality is engineered rather than given, then the interesting question is not whether "the internet" is good or bad for democracy but what a platform must do, at the level of concrete affordances, to produce deliberation rather than merely to host talk. The systematic literature on online deliberation supplies both the norms such a platform should satisfy and a sobering observation about the state of the field: research has concentrated on design intentions and on process, while the measurement of deliberative outcomes remains comparatively underdeveloped [3]. There is, in short, a gap between the platforms scholars study and the platforms scholars would design, and a further gap between describing deliberation and demonstrating that a given design produces it.
This paper addresses the first gap directly and prepares the ground for closing the second. It presents a design analysis of Hoojah, a deployed debate platform whose mechanics were, we argue, chosen to instantiate deliberative norms and whose data model is, as a consequence of those same choices, a machine-readable argumentation structure. Hoojah is a server-rendered application built around a self-referential tree of claims and stance-tagged responses, a three-option vote, a rule that one must vote before one may reply, a phased one-on-one debate format with spectator verdicts, and a secret-ballot regime that suppresses vote breakdowns below a fixed anonymity threshold. Each of these mechanisms, we will show, answers to a specific requirement drawn from deliberative theory, and together they yield a corpus whose structure aligns with the formal apparatus of computational argumentation.
The argument is deliberately two-sided. From the perspective of deliberative democracy, Hoojah is an attempt to compile a set of normative commitments into working software: reciprocity in the form of a reply that presupposes a committed stance, contestation in the form of a bounded adversarial debate, equality in the form of an audience whose verdict rewards argumentative merit, and considered opinion in the form of a ballot whose aggregate is deliberately withheld until it is safe to reveal. From the perspective of computational argumentation, the same design produces data that formal theory can act on: a tree of claims and attacks, every edge of which carries a stance label supplied by the interface rather than inferred by a classifier.
We make three contributions. First (C1), a mechanism-level mapping from deliberative norms to the design of a deployed platform, tagged throughout to distinguish what is built from what is proposed. Second (C2), a formalization of the platform's data model as an argumentation framework, in which top-level claims are argument nodes, stance-labeled replies carry an implicit support or attack polarity, and debates are dialogical sub-games between two committed positions. Third (C3), a built-versus-proposed research roadmap with a concrete measurement plan, so that the design claims of C1 and C2 can be converted into falsifiable hypotheses rather than left as assertions.
The remainder of the paper proceeds as follows. Section 2 reviews the normative core of deliberative democracy, the design-determines-quality thesis in online deliberation, the deliberation systems that precede Hoojah, and the computational-argumentation machinery for reasoning about the result, extracting from each a design requirement. Section 3 builds a two-sided analytical framework, normative and formal, from those requirements. Section 4 walks each of Hoojah's mechanisms and justifies it against the framework, tagging built and proposed inline. Section 5 is a dedicated ledger separating the deployed system from the roadmap. Section 6 turns the framework into an evaluation and research agenda. Section 7 states the limitations, Section 8 concludes.
This review assembles, in four moves, the resources the rest of the paper needs: the norms Hoojah must satisfy, the evidence that design can satisfy them, the systems that have tried, and the formal machinery for reasoning about what results. Each subsection closes by extracting the design requirement against which Hoojah will be tested in Section 4.
The normative baseline is the public sphere as a space of rational-critical debate among equals, in which the coordinating force is argument rather than authority or market power [1]. Habermas's historical account is also a normative one: it furnishes a standard by which actual communicative arrangements can be judged deficient, and it is precisely the standard against which mass-mediated and, later, platform-mediated discourse has been found wanting. For the purposes of a design analysis, the public sphere supplies two enduring commitments: that claims are settled by their merits, and that participants enter as equals whose standing does not depend on their status outside the exchange.
The contemporary theory of deliberative democracy sharpens the first commitment into a norm of reciprocity. On this account, the core of deliberation is that citizens owe one another reasons, and specifically reasons that others could in principle accept, for the positions they advance and the decisions they would impose [4]. Deliberation is not merely talk; it is the mutual exchange of justifications. This gives the analysis its most basic yardstick: a deliberative venue is one whose interactions consist of, and reward, reason-giving rather than mere preference expression or assertion.
A crucial move for this paper is to loosen the association, common in early deliberative theory, between deliberation and the search for consensus. Deliberation need not aim at agreement to be legitimate. A contestatory or discursive conception treats the clash of discourses, and the ongoing challenge of one position by another, as itself a legitimate mode of democratic engagement, distinct from and not reducible to consensus-seeking liberal deliberation [5]. This matters because Hoojah's signature format is adversarial: a structured one-on-one debate between two committed and opposed positions. A theory that licenses only consensus would find such a format suspect; the discursive conception legitimizes it as contestation, provided the contest is conducted through the exchange of reasons.
Finally, the tradition insists that deliberative quality can be produced by procedure. The canonical demonstration is Deliberative Polling, in which ordinary participants, placed in a designed setting with balanced information and structured discussion, arrive at more informed and more considered opinions than they held before [6]. The lesson for platform design is that considered opinion is an achievable output of a well-constructed process, not a trait participants must already possess. Deliberative quality is, in this sense, an engineering target.
It is worth dwelling on why the contestatory turn matters so much for a design analysis of an explicitly adversarial platform. A theory that treats consensus as the telos of deliberation measures every exchange by its distance from agreement, and on that measure a debate that ends in a decided verdict rather than a shared position looks like a deliberative failure. The discursive conception refuses this measure. It holds that the ongoing, structured contest of positions is not a way-station on the road to consensus but a legitimate standing condition of democratic life, in which the point is that positions are tested against their strongest opponents rather than that they are dissolved into a common view [5]. This reframing is what allows Hoojah's one-on-one debate to be read as deliberation at all rather than as its opposite, and it does so without abandoning the reason-giving norm, because the contest the discursive conception licenses is a contest of reasons, not of force or volume.
The requirements extracted here are: interactions should consist of and reward reason-giving (from reciprocity); an adversarial, contestatory format is a legitimate deliberative form (from discursive democracy); and the design should aim to elicit considered rather than merely expressed opinion (from Deliberative Polling).
The linchpin of this paper's method is the finding that the deliberative quality of an online forum is a product of design and choice, not an inherent feature of the technology [2]. The same underlying communication medium can host a shouting match or a structured exchange depending on how the forum is configured, moderated, and framed. This licenses the central analytic stance of the paper: Hoojah is to be read as a set of deliberate design choices, each of which can be evaluated for the deliberative behavior it is likely to produce.
Locating a design's contribution requires a map of what deliberation research measures. The systematic review of the field organizes it into three layers: the design of a deliberative arrangement, the communicative process it produces, and the results or outcomes it achieves; and it observes that the results layer is comparatively under-studied [3]. This three-layer scheme recurs throughout the present paper. It lets us say precisely what a design analysis can and cannot establish on its own: it can characterize the design and predict process effects, but claims about results must await the measurement program of Section 6.
Complementary work supplies the operational vocabulary for such measurement. An early and influential treatment sets out hypotheses, variables, and methods for assessing whether online forums actually deliver deliberation, including indicators of reciprocity, reflexivity, and the diversity of participation [7]. These variables become, in Section 6, the concrete quantities a study of Hoojah would compute over its debate transcripts and threads.
Two further findings shape the requirements. First, civility must be distinguished from politeness. Robust, even heated disagreement can be entirely healthy democratically; what damages deliberation is incivility in the sense of denying others their standing, not the mere presence of passion or force [8]. A platform designed for vigorous contestation must therefore be careful not to engineer away the heat along with the harm. Second, conversational failure is forecastable from its early moments: pragmatic signals present at the start of an exchange predict whether it will later derail into hostility [9], and derailment can be modeled as an emergent property that becomes increasingly predictable as a conversation unfolds [10]. The design implication is that the most effective point of intervention is the format and the opening, not post-hoc moderation of a thread that has already collapsed.
The requirements extracted here are: the design is the independent variable and must be analyzed as such; deliberation should be evaluated with explicit process variables; the format should permit heat while protecting standing; and structure should be imposed early, at the format level and at the start of an exchange, to forestall derailment.
Hoojah is not the first system to try to engineer deliberation, and two precedents are especially instructive. ConsiderIt structures opinion formation around pro and con points that participants create, adopt from others, and arrange into a personal pro/con list on their way to a stance on an issue [11]. Its central lesson is that interface structure shapes deliberative behavior: by making people build and weigh reasons rather than simply react, it elicits more reflective engagement, and by letting them adopt others' points it surfaces shared considerations across a divide. What it demonstrates for our purposes is that structured stance-taking, with reasons attached, is both feasible and behaviorally consequential in a deployed system.
The MIT Deliberatorium attacks a different problem: the incoherence and unmanageable volume of flat, large-scale discussion. It combines a formal argument-map, in which contributions are posted as issues, ideas, and arguments in a logical tree, with attention-mediation metrics that direct participant effort to where it is most needed, so that mass deliberation becomes tractable rather than collapsing under its own scale [12]. Its lesson is that structure plus some principled surfacing of contributions is a mechanism for scaling deliberation, and a remedy for the specific pathology of the flat thread, in which every contribution sits in an undifferentiated pile.
Both systems also reveal what remains open. Argument-mapping and structured pro/con interfaces impose a cognitive and interactional cost that can depress adoption and sustained use; and neither format is built around the sustained adversarial energy of a two-person contest, which is a large part of what makes debate engaging and what channels combative motivation into a productive form. Hoojah targets exactly these openings: it preserves structure but organizes it around a self-referential tree that doubles as an ordinary social feed, and it adds a bounded adversarial debate format that gives contestatory energy a designed container. (A fuller four-system comparison belongs to a companion analysis; here the precedents serve only to locate Hoojah's design choices.)
The requirement extracted here is: retain the structure and scale-management of argument-mapping systems while lowering their adoption cost and providing a designed home for adversarial contestation.
The second theoretical register is formal and computational, and it is what makes the paper's C2 contribution possible. At the level of a single argument, the classic model decomposes an argument into a claim supported by data, licensed by a warrant, and qualified and rebutted under conditions, giving a vocabulary for what a well-formed argument contains beyond its bare conclusion [13]. At the level of reasoning patterns, argumentation schemes catalogue stereotypical forms of defeasible inference, each paired with a set of critical questions that an interlocutor may pose to test it, in a notation designed to be computationally tractable [14]. Together these supply the micro-structure a debate turn could be scaffolded toward, and the critical-question discipline a phased debate protocol can approximate.
At the level of whole debates, abstract argumentation frameworks provide the decisive formal tool. An argumentation framework is a set of arguments together with a binary attack relation between them, and its acceptability semantics determine which sets of arguments can be collectively defended against all attacks, abstracting away entirely from the internal content of the arguments [15]. This is precisely the abstraction a platform of competing claims invites: if claims are nodes and opposing replies are attacks, then the question of which positions are collectively defensible becomes a formal, computable one. The broader landscape of these techniques, from formalisms to systems, is surveyed comprehensively and situates any computational-argumentation component within an established field [16].
The empirical arm of computational argumentation is argument mining: the automatic extraction of argument structure from natural-language text. The definitive survey frames the task as recovering argument components and the relations between them from unstructured prose, a pipeline whose hardest stages are identifying where arguments are and how they relate [17]. Concrete end-to-end methods exist for parsing argumentation structure, for example identifying components and their support and attack relations in persuasive essays, together with annotated corpora to train and evaluate them [18]. A closely related task is stance detection: classifying whether a text is for or against a target. The problem has been studied from ideological online debates using sentiment and arguing-expression features [19], on two-sided debate posts with dedicated stance classifiers [20], and formalized as a community-standard shared task with a fixed favor/against/none label set and evaluation setup [21]. Corpus infrastructure for this line of work includes large annotated collections of online debate posts labeled for agreement, disagreement, and related properties [22], and even modest recent systems demonstrate for/against stance classification over online debates using lightweight features [23].
It is worth being explicit about why capturing structure at the interface is not a mere convenience but a change in kind. Argument mining is difficult precisely because prose does not wear its structure on its surface: the boundaries of an argumentative unit, the identity of the claim it is making, and above all the relation (support or attack) it bears to a neighbouring unit must all be recovered by models that are, on the hardest of these sub-tasks, still far from reliable [17], [18]. Stance detection isolates one of these sub-tasks and formalizes it as a supervised classification problem, which presupposes a labeled training set that someone had to annotate by hand [21], [20]. Every one of these difficulties is a difficulty of recovering structure that the author possessed at the moment of writing but that the medium discarded. A medium that instead retained the structure, that recorded the author's stance and the target of their reply as first-class data, would not make the mining task easier so much as make large parts of it unnecessary.
The requirement extracted here is the one that binds the two registers together: the single most valuable thing a platform can do for computational argumentation is to capture stance and reply-structure explicitly at the interface, because doing so removes the hardest and most error-prone steps of the mining pipeline. A reply whose stance is recorded when it is written does not need to have its stance inferred later; an attack edge that the interface encodes does not need to be recovered by a parser. If a platform's ordinary operation produces exactly the labels that argument mining struggles to predict, then its data are not merely minable but pre-annotated by construction. Section 3 turns this observation into a formal claim about Hoojah's data model.
The literature yields two clusters of requirements, one normative and one formal, and this section assembles them into the two-sided analytical lens that Section 4 applies to Hoojah. The move here is the bridge from problem to solution: Section 2's requirements are converted into an explicit framework with a single testable claim.
We distill five requirements from Sections 2.1 and 2.2. They are stated as design targets so that each of Hoojah's mechanisms can be assessed against them.
R1, Reason-giving. Interactions should consist of, elicit, and reward the exchange of justifications, not the bare expression or aggregation of preferences. This is the operationalization of reciprocity: participants owe one another reasons others could accept [4]. A design satisfies R1 to the extent that advancing or opposing a position on the platform naturally entails offering a reason for it.
R2, Structured contestation. The design should provide a legitimate, bounded container for adversarial exchange between opposed positions, treating the clash of views as a productive democratic form rather than a failure to be suppressed [5]. R2 is satisfied by a format that pairs committed opponents and disciplines their exchange, as opposed to an undifferentiated free-for-all or a consensus-only frame.
R3, Equality of standing. Participants should be treated as equals whose contributions are weighed on their merits, independent of status, reach, or popularity [1]. In a networked setting, R3 bears especially on how outcomes are decided: a mechanism that rewards argumentative merit rather than social diffusion or follower count is the design expression of equal standing.
R4, Considered rather than performed opinion. The design should elicit reflective, considered judgments and avoid conditions that turn opinion into public performance or that let early signals stampede later ones [6]. The empirical warning attached to R4 is specific: visible running tallies induce herding, because a single early positive signal measurably raises the probability and magnitude of subsequent positive signals [24]. A design satisfies R4 partly by what it withholds, namely premature aggregate feedback that would substitute social proof for individual judgment.
R5, Early structure against derailment. Because conversational failure is forecastable from its opening moves [9] and derailment is an emergent property that can be tracked as an exchange develops [10], structure should be imposed early: at the level of format, and at the very start of an exchange, rather than only through after-the-fact moderation. R5 is satisfied by opening scaffolds, turn discipline, and phase structure that shape an exchange before it can curdle.
These five requirements are not independent of one another, and Section 4 will show that several of Hoojah's mechanisms answer more than one at once, which is itself a sign of a coherent design rather than a bag of features.
The second component of the framework reads Hoojah's data model through computational argumentation. The claim is structural, not metaphorical: the platform's ordinary operation produces a graph that the formal apparatus of Section 2.4 can act on directly.
Consider the platform's core object, the self-referential tree of claims and replies. A top-level claim (a hujah with no parent) is naturally read as an argument node: a proposition placed into contention. A reply is a child argument node attached to its parent. What distinguishes Hoojah from an ordinary comment tree is that every reply carries a stance, and it does so not by later inference but by construction, because the platform requires a user to have voted on a claim before replying to it. The stance recorded on that prior vote, agree or disagree, supplies the polarity of the reply's relation to its parent: a reply authored from a disagreeing stance is naturally read as an attack in the sense of an abstract argumentation framework, and a reply authored from an agreeing stance as support [15]. The interface thus hands the formal reading its edges and their signs.
This is the point at which the two registers meet. In the argument-mining pipeline, recovering the reply's stance and the polarity of its relation to its parent is among the hardest steps, the very thing that stance-detection research works to predict [21]. On Hoojah that label is not predicted; it is an explicit, interface-supplied annotation captured at authoring time. The vote-to-respond gate is, from the formal point of view, a mechanism that converts an inference problem into a lookup. The graph that results is a labeled, signed structure over which Dung-style acceptability semantics can be computed to ask which positions in a cluster of competing claims are collectively defensible [15], and which the argument-mining methods of Section 2.4 can parse for finer component structure with their hardest sub-task already solved [17], [18].
One clarification guards against overstatement. The polarity supplied by the vote-to-respond gate records the author's stance toward the parent claim, which is a fact about the author's position rather than a proof of the logical relation between the reply's content and the parent's. A disagreeing author might, in a given turn, concede a point before mounting their attack elsewhere. The formal reading of Section 3.2 therefore treats the stance label as a strong and interface-guaranteed prior on the reply's polarity rather than as an infallible edge sign, and Section 7 records this as a limitation that the argument-mining program of Section 6 is designed to test. The important point for the framework is comparative: even as a prior, an interface-supplied stance label is dramatically more reliable than the predicted label that stance detection must otherwise supply [21], so the graph Hoojah produces starts from a far better position than one recovered from unstructured prose.
The debate object admits a second, complementary formal reading. A one-on-one debate is a dialogical sub-game between two committed and opposed positions, conducted through strictly alternating turns organized into named phases. This is the level at which the micro-structure of argument becomes relevant: each turn is a candidate site for the claim-data-warrant analysis [13], and the phase discipline of a debate approximates the critical-question structure of argumentation schemes, in which each move invites specific challenges from the opposing side [14]. Where the tree gives a static graph of claims and attacks, the debate gives a dynamic, turn-structured dialogue between two nodes of that graph.
The framework's testable claim can now be stated precisely. A platform that satisfies R1 through R5 and whose data natively instantiate the argumentation graph of Section 3.2 is simultaneously two things: a deliberation venue, whose process quality can be measured with the instruments of Section 6, and a research instrument, whose output is a pre-annotated argumentation corpus. The remainder of the paper defends the antecedent (that Hoojah's mechanisms satisfy R1 through R5 and instantiate the graph) and then, in Section 6, specifies how the consequent could be measured.
This is the core of the paper. We walk each of Hoojah's principal mechanisms, describe it exactly as it exists in the system, and justify it against the framework of Section 3. Every mechanism is tagged [built] or [proposed] inline, because the honesty of that distinction is central to the paper's claim.
[built]Hoojah's foundational object is a self-referential tree of claims. A hujah is a single node in a tree that references itself: each node optionally belongs to a parent hujah and has many children, so a top-level claim is a node with no parent and every reply or counter-argument is a child of the node it answers [built]. This is the same recursive structure that argumentation theory posits for a body of contested claims, and it is what allows the formal reading of Section 3.2 to apply without translation: the tree is the argument graph, not a lossy rendering of one.
Two details of the built implementation matter for the deliberative reading. First, a minimum length is enforced on top-level claims only: a top-level hujah must clear a floor on its body length, while replies are left unconstrained [built]. The design intent is legible against R1: a claim entering contention is asked to be substantive enough to be argued with, while a reply, which already sits in an argumentative context and inherits it, is free to be as short as the point requires. Second, the tree carries the ordinary affordances of a social medium, with human-readable slugs derived from a claim's opening words and inline parsing of mentions and hashtags [built]. This is the deliberate answer to the adoption cost that burdened earlier argument-mapping systems (Section 2.3): the structure that makes the data an argument graph is delivered through an interface that behaves like a familiar feed rather than a formal modeling tool.
The justification against the formal axis is direct. Because the tree is self-referential and every node records its author and its relation to its parent, the platform's normal operation continuously produces the node-and-edge structure that Dung-style semantics require [15], with the reply text available for the finer component analysis of argument mining [18], [17]. The tree is the substrate; the stance labels, discussed next, are what make its edges signed.
[built]Every hujah can be voted on with exactly one of three stances, agree, neutral, or disagree, drawn from a closed and single-sourced stance domain [built]. The closure is deliberate and consequential. By fixing the stance domain to three values, the design turns each vote into a categorical annotation over a known label set, which is exactly the shape that stance-detection research treats as its target [21]. The neutral option is not a throwaway; it lets a participant register engagement without a directional commitment, which both respects considered opinion (a participant genuinely undecided is not forced to feign a side) and preserves the integrity of the agree/disagree signal that the formal axis relies on.
The vote record is, as a matter of the built data model, append-only: a user's stance changes are stored as a history, with the last entry taken as the current stance [built]. This began as an implementation detail of a legacy storage choice, but it has a significant and, we argue, valuable deliberative consequence. It means the platform preserves, for every voter on every claim, the full trajectory of their opinion over time. This is precisely the quantity that deliberative-polling logic wants to observe: the movement of considered opinion as participants encounter argument [6]. A design built to study deliberation would deliberately instrument opinion change; Hoojah records it as a side effect of how votes are stored, and Section 6 treats this longitudinal record as a research asset. We flag that the collapse of this append-only array into a single scalar, together with the unification of the several stance-bearing columns into one enumerated type, is a proposed data-model cleanup [proposed] (the "stance-domain unification" backlog item); it would tidy the model without changing the mechanic, and would in fact require care to preserve the opinion-change history that the current representation captures for free.
Layered on the vote is conviction: a voter may mark a vote as a conviction, which is a costly and irrevocable commitment rather than a weight [built]. A conviction vote locks the stance permanently, so that the voter forfeits the ability to change their position on that claim, and only an aggregate count of convictions is retained. The design reading is twofold. Against R1 and R4, conviction is a considered-opinion signal with skin in the game: it distinguishes a position a participant is merely willing to register from one they are willing to be held to, and it does so through irrevocability, a genuine cost, rather than through a cheap emphatic display. Against the secret-ballot commitment discussed next, conviction is carefully constructed to add no attribution surface, because only a count, never a per-voter conviction record exposed to others, is kept.
[built]Hoojah treats voting as an effectively secret ballot, and the built system enforces this through several coordinated choices. There is no serialization of individual votes: the platform has no vote serializer and never exposes a vote object, and a user can retrieve only their own current stance, never anyone else's [built]. Vote filtering for blocked users is deliberately omitted precisely because votes carry no attribution and therefore present no vector to exploit [built].
The load-bearing case concerns vote notifications, and here the paper must be precise because the platform's own older roadmap describes a flaw that the shipped code has closed. When a claim receives its first vote, the author is notified that a vote landed, but the notification deliberately carries no identifier of the voter: the subject-user field that would name the voter is absent by construction, and the serializer that renders notifications emits a username only when that field is present, so the API never hands a claim's owner the identity of who voted [built]. A one-off backfill migration additionally nulled the voter identifier on all pre-existing vote notifications, closing the historical residue as well as the live path [built]. This is the resolution of what an earlier roadmap, dated in the platform's development history, had flagged as a live de-anonymization vector; we treat that roadmap warning in Section 5 as a historiographic note, because the authoritative artifact is the shipped code, and the shipped code closed the vector.
Beyond identity, the design also suppresses premature aggregate feedback. The per-stance breakdown of a claim's votes (how many agreed, how many disagreed) is withheld until the total number of votes reaches a fixed anonymity threshold of five, below which only the total count is shown and the breakdown is null, and this suppression is applied uniformly, including to the claim's own author [built]. The threshold is drawn from a single shared source so that it cannot drift between features. The justification is squarely evidential. Visible running tallies induce herding: a single early signal measurably raises the probability and size of later signals in the same direction, distorting the very quantity a deliberative venue exists to elicit [24]. Suppressing the breakdown below a safe threshold protects considered opinion (R4) by ensuring that early voters form their judgment before any aggregate can stampede it, and by treating the author as an observer too, so that even the person with the most interest in the tally cannot use it to read the room prematurely.
Two residuals must be marked honestly. First, the first-vote notification still reveals, by its timestamp, that some vote landed and when, though never who or which choice; this is an accepted residual inherent to any activity notification [built, accepted]. Second, certain aggregate count columns elsewhere in the system remain unfiltered, and a fully specified end-state for the secret-ballot regime (recorded in the roadmap as a decided but not-yet-implemented option) is outstanding [proposed]. These are genuine limits, not defeats of the design: the central de-anonymization vector is closed, the herding vector is mitigated by k-anonymous suppression, and what remains is a smaller surface tracked for future work.
[built]The mechanism that most tightly binds the normative and formal axes is the rule that one must vote on a claim before one may reply to it. In the built system, authorizing a reply requires that the replying user has already voted on the parent claim (alongside visibility and block checks) [built]. Commitment precedes contestation.
Against the normative axis, the gate is a direct operationalization of R1 and R2. It ensures that no one argues without first taking a position, which enforces the reciprocity structure of deliberation, that one enters the exchange as a committed party owing reasons rather than as a detached heckler [4], and it channels reply energy into structured contestation between stances rather than into undifferentiated commentary [5]. There is a modest friction cost, which Section 6 treats as an empirical question; the design bet is that the cost buys commitment.
Against the formal axis, the gate is what makes the whole enterprise of Section 3.2 work. Because a reply cannot exist without a prior vote, every reply is stance-labeled at authoring time, and the polarity of its relation to the parent is given, not guessed. The consequence for computational argumentation is exact: the stance-detection task that the field formalizes as a prediction problem [21] is, on Hoojah, an explicit stance label supplied by the interface and recorded with the reply. The gate converts stance from an inference problem into a database fact. This single mechanism is the strongest evidence for the paper's thesis that a design built to deliberative norms can be, at the same time and without additional instrumentation, a research instrument for argumentation.
It is worth noticing that the gate achieves this dual payoff with a single rule rather than two. A design that wanted both the deliberative benefit (no arguing without commitment) and the formal benefit (every reply carries a stance label) might naively implement them as separate features: a commitment prompt on one hand, a stance annotation on the other. Hoojah collapses them. The commitment the deliberative norm demands is exactly the annotation the formal axis needs, because a stance is at once a substantive act of taking a position and a categorical label over the closed domain of Section 4.2. This coincidence is not accidental; it is the reason the platform can be a deliberation venue and a research instrument without paying twice, and it is the clearest illustration of the paper's central claim that the two registers, normative and formal, are not merely compatible on Hoojah but are served by the very same mechanisms.
[built]Where the tree hosts many-to-many contestation, the debate format provides the designed container for sustained adversarial exchange that Section 2.3 identified as the opening earlier systems left. A debate is created against a specific claim and anchored to an argument, between a challenger and an opponent who must hold differing stances (the model validates only that the two stances are not identical, not that they are strict opposites, so an agree-versus-neutral pairing is permitted as well as agree-versus-disagree) [built]. It proceeds through a lifecycle of challenge, accept or decline, active exchange, and conclusion [built].
Several built details bear directly on the framework. The exchange is strictly alternating: turns are posted in sequence, and at any moment exactly one participant is the mover, namely whichever party did not author the last turn [built]. The debate is organized into named phases: an opening statement, counter-arguments, responses, and a closing statement, with the phase of each round derived from its position in the debate rather than stored as free-form state [built]. This phase structure is a lightweight dialogical protocol, and we read it in the spirit of the critical-question discipline of argumentation schemes, in which each type of move sets up the specific challenges the opposing move should answer [14]: an opening establishes a position, a counter-argument attacks it, a response defends, and a closing consolidates. The exchange is bounded, with a configurable round limit within a fixed ceiling and a single consensual extension available only at the closing-round boundary, guarded by a row lock so the phase labels cannot be corrupted by concurrent turns [built]. And an idle debate auto-concludes after seven days, with the idle clock tracking the last turn rather than the debate's creation, so that an actively argued debate is never wrongly closed [built].
The justification is layered. Against R2, the format is the platform's clearest instantiation of contestatory deliberation: two committed, opposed positions, disciplined into a fair alternating exchange [5]. Against R5, the opening structure and phase discipline are exactly the early, format-level intervention that derailment research prescribes: rather than waiting for a thread to turn toxic and moderating it after the fact, the debate imposes structure at the start, when the trajectory of an exchange is still malleable and its eventual failure is, in principle, already forecastable [9], [10]. And the format is deliberately designed to permit heat while protecting standing, consistent with the finding that vigorous disagreement is democratically healthy so long as it does not deny participants their standing [8]: the alternation, bounds, and phases constrain how the disagreement is conducted without dampening its force. The precedent systems inform this design: the debate carries forward the structured stance-taking of ConsiderIt [11] and the structure-for-scale logic of the Deliberatorium [12], while adding the adversarial container neither provided. Read formally, a debate is the dialogical sub-game of Section 3.2, a turn-structured dialogue between two nodes of the argument graph whose transcript is a candidate for the claim-data-warrant analysis of individual moves [13].
[built]A debate is judged by its audience. Any non-participant who can see a concluded debate may cast exactly one immutable verdict for the challenger, the opponent, or a draw, and the winner is derived on read from the tally, with any tie, including an empty tally or a draw plurality, resolving to a draw and only a unique maximum crowning a side [built]. Verdicts are one-per-spectator and cannot be changed once cast [built].
The design reading is against R3. The spectator verdict is an equality-of-standing mechanism: it decides the outcome of a contest by the judgment of the audience on the merits of the arguments, not by which participant has more followers, more reach, or a louder network. In a platform whose surrounding medium decides prominence by diffusion, a verdict that rewards argumentative merit is the concrete design expression of the public sphere's commitment that claims are settled among equals by the force of the better argument [1]. This is the point at which Hoojah's design most directly inverts the engagement-ranking pathology named in the introduction. Where a diffusion-driven feed lets the outcome of a disagreement be decided by which party can mobilize the larger or more active network, the verdict relocates the deciding judgment to a body of spectators each of whom gets exactly one immutable say, weighted by nothing but the fact of having watched the exchange. The verdict is thus not merely a scoring convenience; it is the mechanism through which the platform refuses to let social reach substitute for argumentative merit, which is the substance of what equal standing demands in a networked setting. The immutability and one-per-spectator constraints protect that judgment from the same herding and manipulation pressures the secret ballot guards against, drawing on the same evidence that visible early signals distort later ones [24], and the conservative tie-to-draw rule refuses to manufacture a winner where the audience did not clearly find one, which is itself a considered-opinion commitment (R4): the design would rather record honest indecision than fabricate a decisive result. We mark one proposed extension: a debate-won form of recognition (a badge) awaits a finalized verdict-tally rule and is therefore not yet built [proposed], a deliberate hold rather than an omission, since awarding recognition on an unsettled rule would bake a contestable judgment into a permanent record.
The credibility of a forward-looking design paper rests on an honest separation of what is deployed from what is merely planned. This section is that ledger. It is organized as a status audit and closes by classifying the proposed items by their research significance, which is what makes the roadmap relevant to the paper's contribution rather than a mere feature list. In the language of the field's three-layer scheme, the built column is the design and much of its process apparatus, while the proposed column is largely what remains before the results layer can be measured [3].
Built (in production). The following are present and working in the source tree, on a test suite reporting 536 examples with zero failures. The full claim-vote-argue loop: the self-referential hujah tree, three-stance voting with conviction, and the vote-to-respond gate. The complete debate lifecycle: challenge, accept and decline, strictly alternating phased turns, bounded rounds with a single consensual extension, real-time delivery of turns over the application's websocket surface, spectator verdicts with compute-on-read tallies, and idle timeout auto-conclusion. The secret-ballot hardening: no vote serialization, the de-anonymization vector via vote notifications closed and backfilled, and k-anonymous suppression of vote breakdowns below a total of five, applied uniformly including to the author. The safety and privacy model: flag-and-moderation with staff-only visibility of removed content, a follow graph with request-and-approve for private accounts, and a bidirectional block model enforced at the authorization layer. These are the mechanisms Section 4 justified; they are shipped.
Proposed (roadmap, not in code). The following are named in plans and not yet implemented: identity verification; the stance-domain unification that would collapse the append-only vote array to a scalar and unify the several stance columns under one enumerated type; analytics rollups for trends over time, reach, and impressions; bookmarks; debate-won and vote-milestone recognition badges; full visibility-and-block parity for the JSON API consumed by future native clients; native clients themselves; the removal of the remaining unfiltered aggregate count columns and the fully specified secret-ballot end-state; and the concurrency fix for a first-vote uniqueness race, in which two simultaneous first votes can both create a vote row before either sees the other, to be closed with a unique index and a rescue. Each is genuinely absent from the running system and is presented as such.
A historiographic note. The platform's own roadmap, dated during development, describes the vote-notification de-anonymization vector as live, stating that a vote notification carries the voter's identity and that an owner therefore already learns who first voted on their claim. That description is accurate to a pre-hardening state of the code and is now obsolete: the branch that shipped the secret-ballot work removed the voter identifier from the notification and backfilled the historical rows. We follow the dossier's methodological rule that the authoritative artifact is the code, not the prose documentation, and prose documentation can lag the code it describes. The vector is resolved; the roadmap's warning survives only as a record of the problem the design solved.
Classifying proposed items by research significance. Not all proposed work matters equally to the research program of Section 6, and sorting it by what it unblocks is more useful than sorting it by feature area. Some items unblock measurement: the analytics rollups, in particular, are what would turn the platform's process data into the longitudinal series a results-layer evaluation needs, and their absence is the main reason Section 6 is an agenda rather than a results section. Some items unblock formalization: the stance-domain unification would give the argumentation graph of Section 3.2 a single clean stance type across claims, votes, and debates, simplifying any computational analysis over it, though the mechanic it tidies is already fully functional. And some items unblock trust: identity verification bears on the interpretation of every deliberative claim in the paper, since the strength of an equal-standing argument (R3) depends on who the equals are, and it is properly held back as a distinct future concern with its own design surface. The remaining items (bookmarks, badges, native clients, the count-column residuals, the race fix) are refinements that do not gate the core thesis.
A design analysis establishes that Hoojah is built to deliberative norms and that its data instantiate an argumentation graph. It does not, on its own, establish that the platform produces deliberation, which is a results-layer claim the field itself flags as the hardest to make [3]. This section turns the framework into falsifiable next steps along three threads: measuring deliberative quality, exploiting the data computationally, and confronting the design tensions the analysis exposes. This is what makes the paper forward-looking research rather than a system description: it specifies what would confirm or refute the design's claims.
The first thread applies established instruments to Hoojah's output. The Discourse Quality Index provides a Habermas-derived, behaviorally coded measure of deliberation, scoring contributions on dimensions such as the level of justification offered, whether reasons appeal to the common good, and respect for others' demands [25]. Debate transcripts, with their clean turn boundaries and phase labels, are unusually well suited to such coding, because the unit of analysis (the turn) is given by the platform rather than imposed by the coder. Stromer-Galley's content-analytic coding scheme, built for online and face-to-face groups, complements the DQI with categories tuned to the messy realities of mediated discussion [26], and Janssen and Kies's variables for online forums supply the indicators (reciprocity, reflexivity, diversity of participation) for the tree layer, where exchanges are many-to-many rather than dyadic [7].
With instruments in hand, the design claims of Section 4 become hypotheses. The phase-structure hypothesis: phase-structured debates score higher on reason-giving and reciprocity than matched flat threads on the same claims, which would confirm that the R5 opening-and-phase intervention produces deliberative process gains and not merely a tidier interface. The commitment hypothesis: conviction votes correlate with higher argumentative engagement (more and longer replies, more debate participation) than ordinary votes, which would confirm that the R4 costly-signal design selects for considered rather than performed opinion. The gate hypothesis: threads whose replies passed through the vote-to-respond gate exhibit higher reciprocity than comparable ungated comment threads elsewhere, which would confirm the R1 justification of Section 4.4. Each hypothesis is falsifiable with the coding instruments above and a suitable comparison set, and each maps to a specific requirement, so a null result would localize which part of the design failed to deliver.
Two features of Hoojah's built data model make this program unusually tractable, and both are worth naming because they distinguish a study of Hoojah from a study of an ordinary forum. The first is that the unit of analysis is given rather than imposed. Coding schemes such as the DQI and Stromer-Galley's must first segment a stream of talk into codeable units, a step that introduces coder disagreement before any substantive coding begins [25], [26]; on Hoojah the turn and the reply are already discrete, authored, and stance-labeled objects, so the segmentation step largely disappears and the reliability of the coding rises accordingly. The second is the longitudinal opinion record of Section 4.2. Because the vote history is append-only, a study can observe not only the final distribution of stances on a claim but the trajectory by which it was reached, which is exactly the before-and-after movement that deliberative-polling designs go to great lengths to instrument [6]. A natural extension of the commitment hypothesis follows: stance changes that occur after a participant has engaged with an opposing argument (a reply or a debate) should, if the platform is producing considered opinion, be more stable and more often toward the better-argued position than changes that occur without such engagement. This is a results-layer claim in the strict sense of the field's three-layer scheme [3], and it is answerable precisely because the platform records the process that produced the result.
The second thread exploits the formal axis. Hoojah's stance-labeled replies are, as Section 4.4 argued, free supervision for stance detection: a stream of text paired with ground-truth for/against labels of exactly the kind that stance-detection research must otherwise annotate by hand or predict [21], [19], [20]. This inverts the usual data economy of the field. Where prior work built classifiers to recover stance from debate text, a Hoojah corpus supplies the labels for free and lets the research question move downstream: how well do models trained on interface-supplied stance labels generalize, and what does a large, natively labeled corpus reveal that hand-annotated collections could not.
The debate transcripts constitute an argument-mining corpus that complements existing collections such as the Internet Argument Corpus [22], with the advantage that its top-level structure (who argued, on what side, in what phase) is given rather than annotated. Over this corpus, the parsing methods of the field can recover the finer component structure within turns, the claims, premises, and their relations [18], [17], with their hardest sub-task (stance and reply-polarity) already solved by the interface. Because the phase of every turn is recorded, the corpus also supports a question the flat collections cannot easily pose: whether the internal argument structure of a turn varies systematically by its dialogical role, so that opening statements, counter-arguments, responses, and closings exhibit distinguishable patterns of claim-and-premise organization. That the phase discipline of Section 4.5 was designed to approximate the critical-question structure of argumentation schemes [14] makes this more than a curiosity: a finding that counter-argument turns disproportionately attack the premises rather than the claims of the opening, for instance, would be evidence that the phase scaffold is inducing the reasoning behaviour it was meant to induce.
The tree as a whole invites the abstract-argumentation analysis of Section 3.2: computing Dung-style acceptability over a cluster of competing claims and their signed reply-edges yields a "collectively defensible positions" analytic, a formal readout of which stances in a contested area survive the attacks mounted against them [15]. This is a genuinely different kind of aggregation from a vote count. A vote count reports how many hold a position; a Dung-style acceptability computation reports which positions can be defended against every attack raised against them, which is closer to what deliberation is supposed to produce and further from the popularity contest that engagement-ranked feeds reward. The situating survey of computational argumentation locates such analyses within a mature body of formalisms and algorithms that could be applied to the Hoojah graph without invention [16]. Finally, the derailment-forecasting line of work [9], [10] suggests a concrete [proposed] feature grounded in R5: a live nudge that reads the developing dynamics of an active debate and prompts participants before an exchange collapses, converting the retrospective finding that failure is forecastable into a prospective intervention. Because Hoojah's debates are already segmented into turns and phases and delivered in real time over the platform's websocket surface, the signal such a forecaster would consume is available as the debate unfolds, which is precisely the condition the forecasting work assumes [10].
The third thread is honest about the trade-offs the analysis surfaces, because a design that claimed to have resolved every tension would be less credible, not more. Three tensions stand out.
Friction versus participation. The vote-to-respond gate (Section 4.4) and the conviction lock (Section 4.2) buy commitment with friction, and friction can depress participation. Whether the commitment gained is worth the participation lost is an empirical question that Section 6.1's engagement measures, run against the friction points, are designed to answer.
Suppression versus transparency. The k-anonymous tally suppression (Section 4.3) protects considered opinion by withholding information that a fully transparent design would show, and there is a genuine value conflict between the anti-herding rationale [24] and a competing norm of openness about a public vote. The chosen threshold is a defensible point on that spectrum, not a dissolution of the conflict, and the right threshold is itself an empirical target.
Contestation versus consensus. The deepest tension is theoretical. Hoojah's debate-and-verdict format is contestatory in the discursive sense [5], deciding contests by adversarial argument and audience judgment, while its secret ballot and k-anonymous tallies are consensus-tempering, refusing the social-proof cascades that manufacture false agreement and preserving the space for reciprocal reason-giving [4]. Rather than choosing a side in the long debate between contestatory and consensus-oriented deliberation, Hoojah operationalizes both poles at once: its verdicts are contestatory, its tallies consensus-tempering. Whether a single platform can coherently serve both, or whether the two commitments interfere in practice, is perhaps the most interesting question the design raises, and it is answerable only by the measurement program above.
The paper's claims are bounded in five ways, and each limitation is paired with the step in Section 6 that would address it.
First, this is a design analysis, not yet an outcome evaluation. We have argued that Hoojah is built to deliberative norms and instantiates an argumentation graph; we have not shown that it produces measurably better deliberation, which is precisely the results-layer claim the field identifies as under-evidenced [3]. This limitation is the reason Section 6.1 exists, and it will bind until that program is run.
Second, the analysis concerns a single platform deployed in one national context, which limits the generalizability of any behavioral claim. We frame this less as a defect than as future comparative purchase: a specific deployment context is an asset for the cross-platform and cross-cultural comparative studies that Section 6 anticipates, provided its specificity is treated as a variable rather than ignored.
Third, the formal mapping of Section 3.2 is a proposed reading, not a validated one. Reading agreeing replies as support and disagreeing replies as attack in the sense of an abstract argumentation framework is a principled approximation, but it is an approximation: the stance label records the author's position on the parent, which is not identical to the logical relation between the reply's content and the parent's, and the finer Toulmin micro-structure of any given turn [13] is not yet extracted or validated against the assumed polarity. Section 6.2's argument-mining program is what would test how faithful the approximation is.
Fourth, users of a debate platform are self-selected, and people who choose to join a venue built for structured contestation are not a random sample of the public. Any measured deliberative quality must therefore be read against that selection, and the comparison designs of Section 6.1 must control for it rather than attributing to the design what may be an artifact of who shows up.
Fifth, the proposed features of Section 5 are unimplemented and may change or be cut; nothing in the paper's core thesis depends on them, but the research significance we assigned them in Section 5 is contingent on their eventual form. We have been careful throughout to rest the argument on built mechanisms and to mark proposed ones as such, so that this limitation touches the agenda, not the analysis.
The deficit this paper began from is real and structural: mainstream online discourse is organized in ways that negate the public sphere it was once hoped to enlarge, replacing reasoned exchange among equals with flat threads, engagement ranking, and reaction over reason [1]. The response we have defended is that this deficit is a design outcome and therefore a design problem: deliberative quality is not a fixed property of the medium but a product of choice, and a platform can be built to instantiate deliberative norms directly in its mechanics [2].
Hoojah is our worked case that such compilation is possible. Its self-referential claim tree, three-stance vote, vote-to-respond gate, phased one-on-one debate, spectator verdicts, and k-anonymous secret ballot are not an assortment of features but a coordinated instantiation of five deliberative requirements: reason-giving, structured contestation, equality of standing, considered opinion, and early structure against derailment. And because those same mechanisms record stance and reply-structure at the interface, the platform's ordinary operation produces, by construction, a machine-readable argumentation graph whose hardest annotation is a database fact rather than a prediction. The design is thus, at once, a public sphere in miniature and an argumentation dataset in the making.
What remains is measurement. We have offered a mechanism-level mapping from norms to a deployed design, a formal reading of that design as an argumentation framework, and an honest ledger of what is shipped against what is proposed, together with a falsifiable agenda for turning the design's claims into evidence. The invitation the paper closes with is to run that agenda: to code the transcripts, mine the graph, test the hypotheses, and so establish whether norms compiled into software do, in the event, produce the deliberation they were built to elicit.