NLP & Language Models
87 beliefs (87 IN, 0 OUT)
-
IN
andrew-ng-transfer-learning-next-driver-ml
Andrew Ng at NIPS 2016 predicted transfer learning would be the next major driver of ML commercial success after supervised learning. -
IN
cascade-correlation-constructive-algorithm
Cascade-Correlation (Fahlman & Lebiere, 1991) is a constructive learning algorithm that dynamically adds hidden units during training rather than fixing network architecture in advance. -
IN
convergent-necessities-and-existence-proof-jointly-stranded
ML's two independent sources of mathematical reliability knowledge — convergently discovered necessities (gradient computation, weight sharing, gradient flow) and the SVM existence proof (demonstrating reliable ML is Bayes-optimal, not just achievable) — are jointly stranded by the same pragmatism paradox that enabled their discovery, establishing that ML's reliability knowledge is complete yet completely disconnected from its capability trajectory. -
IN
convexity-determines-both-mathematical-quality-and-economic-fate
Optimization landscape topology appears to influence both a paradigm's theoretical robustness and its economic trajectory — convexity contributes to SVMs' mathematical elegance and guaranteed global optimality but coincides with the scaling barriers that economically strand them, while non-convex minimax landscapes enable GANs' capability but undermine their theoretical guarantees. This suggests a tension where properties like convexity that support mathematical reliability may work against economic favorability, though the evidence from these two cases is insufficient to establish this as a general principle. -
IN
convexity-tragedy-geometry-determines-reliability-economics-anti-correlation
Optimization landscape geometry creates a tragic anti-correlation between mathematical reliability and economic viability — convexity simultaneously produces global optimality guarantees (SVMs' Bayes-optimal reliability), scaling barriers (quadratic complexity), and economic stranding, while non-convexity simultaneously produces theoretical fragility (no guaranteed equilibria), scalability, and economic success — meaning the mathematical property that guarantees reliability is the same property that guarantees economic failure, and this is geometrically determined rather than contingent. -
IN
diagnostic-instruments-confirm-own-obsolescence
ML's diagnostic instruments confirm their own obsolescence from two directions — universal diagnostic capacity (error decomposition into bias, variance, and irreducible noise) is enclosed within the field's epistemic closure and can only confirm the crisis it cannot resolve, while the SVM existence proof that reliable ML is achievable becomes asymptotically irrelevant as capability scaling widens the gap, meaning both the analytical and constructive demonstrations of what ML knows point toward the futility of that knowledge. -
IN
discovered-mechanisms-economically-gated-for-decades
ML's foundational mechanisms appear to be discovered rather than invented, as evidenced by independent convergence across disconnected fields, and these mechanisms are subject to systematic multi-decade adoption latencies — transfer learning (1976 to 2010s) and self-supervised pretraining (1991 to 2018) were both available long before widespread use, suggesting that economic and hardware readiness rather than idea availability is a primary gate on ML progress. This pattern raises the possibility that other already-discovered insights may similarly remain stranded between discovery and deployment until enabling conditions emerge. -
IN
discovered-truths-and-existence-proof-jointly-inert
ML's two independent sources of mathematical reliability knowledge — convergently discovered necessities (backprop, weight sharing, gradient flow) and the SVM existence proof of achievable reliability — are jointly trapped in epistemic inertness, the existence proof absorbed into a state of perfect-knowledge-zero-consequence and the discovered truths trapped in the epistemic fixed point, eliminating both the constructive component ("how to build reliable systems") and the existential component ("that reliable systems are possible") of any reform program. -
IN
discovered-truths-trapped-in-epistemic-fixed-point
The mathematical necessities that pragmatism enabled discovering — gradient computation, weight sharing, gradient flow solutions, each independently validated across disconnected fields — are trapped within the epistemic fixed point that pragmatism simultaneously created, meaning the genuine mathematical truths needed for reliable systems exist within the field's knowledge but cannot escape the self-sustaining, self-amplifying reliability gap that is fully characterized yet structurally irresolvable. -
IN
domain-adaptation-subproblem-transfer-learning
Domain adaptation is a sub-problem of transfer learning focused on when source and target domains differ but the task is the same. -
IN
dual-lock-in-renders-existence-proof-doubly-irrelevant
ML's pragmatism-rooted dual lock-in renders the SVM existence proof of reliable ML irrelevant along every possible dimension of pursuit — economic lock-in ensures no market force will invest in mathematical reliability over scalable capability, epistemic lock-in ensures no intellectual force will recognize the proof's foundational significance, and the widening self-amplifying reliability gap ensures the proof recedes asymptotically from practical reach regardless of recognition or investment. -
IN
existence-proof-absorbed-into-inert-knowledge
The SVM existence proof that reliable ML is mathematically achievable has been absorbed into ML's state of perfect knowledge with zero institutional consequence — rather than serving as a blueprint for reform, the demonstration that theory-practice unity is possible joins the complete corpus of self-knowledge (root cause identified, crisis empirically confirmed, gap self-amplifying) that the field possesses but cannot act upon. -
IN
existence-proof-asymptotically-irrelevant
The SVM existence proof that reliable ML is mathematically achievable becomes asymptotically irrelevant as the self-amplifying reliability gap widens — SVMs demonstrate theory-practice unity is possible, but the gap's self-originating and self-amplifying nature means the distance between achievable and accessible grows without bound, rendering the existence proof increasingly academic with each generation of capability scaling. -
IN
existence-proof-recedes-with-capability-scaling
The distance between achievable and actual reliability grows with capability scaling — SVMs prove reliable ML is mathematically achievable, but scaling simultaneously increases both the potential for harm and the impossibility of accountability, making the existence proof increasingly tantalizing as a demonstration and increasingly irrelevant as a practical guide. -
IN
feature-engineering-canary-for-crisis
The persistence of manual feature engineering is a canary for ML's deeper crisis dynamic — it reflects not just the manifold hypothesis's incompleteness as a practical guide but the broader pattern where pragmatism creates capabilities (deep learning's partial automation of representation) without the theoretical depth to complete them, mirroring the innovation-without-reliability pattern at the methodology level. -
IN
federated-decentralization-tensions-with-data-integrity
Federated learning's privacy-preserving decentralization creates a structural tension with data integrity — distributing training across user devices prevents central data inspection, making federated systems inherently more vulnerable to data poisoning attacks than centralized training, as malicious data injections cannot be detected or filtered by a central authority that never sees the raw data. -
IN
four-major-learning-paradigms
The four major neural network learning paradigms are supervised, unsupervised, reinforcement, and self-supervised learning. -
IN
generative-pretraining-paradigm-unsupervised-then-finetune
The generative pretraining paradigm (train unsupervised, then fine-tune supervised) is the foundation of modern LLMs -
IN
idea-latency-validates-economic-gating
ML ideas exhibit systematic multi-decade adoption latencies — transfer learning (invented 1976, adopted 2010s) and self-supervised pretraining (invented 1991, dominant 2018) were both available for decades before widespread use, providing independent evidence that ML progress is gated by economic and hardware readiness rather than idea availability. -
IN
llm-training-pipeline-stages
Modern LLM training pipeline built on transformers: self-supervised learning → fine-tuning → instruction tuning → RLHF/Constitutional AI -
IN
manifold-geometry-only-surviving-theoretical-anchor
The manifold hypothesis stands out as a relatively robust theoretical anchor in ML — it provides a non-biological foundation spanning the full architecture spectrum, while much of ML's broader theoretical apparatus (generalization theory, paradigm taxonomy, practical-theoretical alignment) remains in a weakened or revisionary state. This makes manifold geometry a comparatively strong candidate for principled reasoning about architecture design, though the overall theoretical landscape's instability means even this foundation should be held with appropriate uncertainty. -
IN
math-determines-failure-mode-economics-determines-timing
Mathematical impossibility results and economic selection pressures play orthogonal roles in paradigm evolution — GAN Nash impossibility (Farnia & Ozdaglar 2020) mathematically necessitated the specific failure mode (training instability, mode collapse) but did not determine adoption or displacement timing, which was governed by scalability economics; mathematics determines HOW paradigms fail while economics determines WHEN. -
IN
mathematical-completeness-counterproductive-for-survival
Mathematical completeness can become counterproductive for paradigm survival in ML — SVMs illustrate how completeness creates its own scaling barriers (three decades of development produced complexity that compounds with problem size), while broader evidence suggests that neither theoretical elegance nor empirical dominance is sufficient to guarantee persistence, complicating the expected value of mathematical rigor. -
IN
mathematical-foundations-economically-stranded
ML's mathematical foundations are economically stranded — convergently discovered as genuine mathematical necessities across independent fields, yet the economic trajectory that governs ML's evolution systematically sustains the misalignment between theory and practice, leaving validated mathematical foundations permanently disconnected from the deployed systems that could benefit from them. -
IN
mathematical-quality-orthogonal-to-evolutionary-success
Mathematical quality alone does not determine paradigm survival in ML when economic selection pressure dominates — SVMs achieved strong theory-practice unity through intellectual selection pressure but face scaling barriers that economically strand their mathematical foundations, while GANs gained unique capabilities through pragmatic selection but inherited fundamental training instability despite sophisticated analytical characterization. This suggests that the type of selection pressure shaping a method is a primary factor in its methodological reliability and evolutionary trajectory, and that validated mathematical foundations can remain permanently disconnected from deployed systems when economic incentives sustain the misalignment. -
IN
ml-acm-2012-subcategories
The ACM 2012 Computing Classification System categorizes Machine Learning under Artificial Intelligence with sub-categories: Supervised, Unsupervised, Reinforcement, Multi-task, and Cross-validation. -
IN
ml-association-rules-no-item-order
Association rule learning discovers relationships between variables in databases (e.g., market basket analysis) but does not consider item order, in contrast with sequence mining -
IN
ml-conceptual-foundations-doubly-unstable
ML's conceptual foundations are doubly unstable — the classical paradigm taxonomy (supervised/unsupervised/RL) is dissolving as modern pipelines combine all three, while even dominant paradigms like GANs and pretrain-finetune prove empirically fragile and transient — suggesting that ML's organizing categories are descriptive conveniences rather than natural kinds. -
IN
ml-federated-learning-decentralized
Federated learning decentralizes training across user devices, preserving privacy by not sending raw data to a central server (e.g., Google Gboard) -
IN
ml-foundationally-bounded-and-task-relative
ML is foundationally bounded to task-relative operation — Mitchell's definition structurally requires domain-specific choices (task T, experience E, measure P) and the No Free Lunch theorem mathematically proves no universal escape exists, jointly establishing task-relativity as an invariant of the field, not a limitation to overcome. -
IN
ml-no-free-lunch-theorem
The No Free Lunch theorem states that no single machine learning algorithm works best for all problems -
IN
ml-pragmatism-principle-triply-validated
ML's pragmatic-over-rigorous character is triply validated across independent domains — backpropagation succeeds precisely when its mathematical prerequisites are violated, mathematical completeness is counterproductive for paradigm survival (SVMs), and NLP's paradigm succession tracks hardware scalability not theoretical adequacy — establishing pragmatic scalability as ML's dominant evolutionary law. -
IN
ml-self-supervised-subset-unsupervised
Self-supervised learning is a special case (subset) of unsupervised learning that generates supervisory signals from the data itself, not a separate paradigm -
IN
ml-three-classical-paradigms
The three classical machine learning paradigms are supervised learning (labelled data), unsupervised learning (no labels), and reinforcement learning (reward signal) -
IN
modern-pipelines-dissolve-classical-paradigm-taxonomy
Modern LLM training pipelines dissolve the classical three-paradigm taxonomy — self-supervised pretraining blurs the supervised/unsupervised boundary (its taxonomic status is actively debated), and the full pipeline synthesizes all three paradigms sequentially, suggesting the taxonomy was always a pedagogical convenience rather than a natural partition of learning. -
IN
modern-pretraining-dominant-but-transient
Modern pretraining as industrial-scale transfer learning is simultaneously the most successful ML methodology and the most likely to be displaced — its dominance rests on empirically fragile foundations (pretraining can hurt), and the broader pattern of paradigm succession (GANs→diffusion) suggests today's self-supervised pipelines are transient. -
IN
more-training-data-often-beats-algorithm-tuning
Given fixed resources, collecting more or better training data often outperforms tuning algorithms in supervised learning -
IN
mrdtl-extends-decision-trees-to-relational-databases
Multi-relational Decision Tree Learning (MRDTL) extends decision trees to relational databases using selection graphs as decision nodes and tuple ID propagation to reduce redundant operations. -
IN
muhlhoff-five-types-machinic-capture
Mühlhoff identified five types of 'machinic capture' for training data collection: gamification, trapping/tracking (CAPTCHAs), social motivation exploitation, information mining, and clickwork. -
IN
negative-transfer-performance-degradation
Negative transfer occurs when source domain knowledge hurts target task performance; transferring from an unrelated domain can degrade results. -
IN
nlp-most-distant-from-reliable-ml
NLP represents the ML domain most distant from reliable ML — it is simultaneously the domain where crisis is most advanced and least remediable (most capable methods are least interpretable, most data-hungry, and most susceptible to hallucination) AND where the SVM existence proof is most irrelevant (the distance between achievable and actual reliability grows most rapidly in the domain where capability scaling is most extreme). -
IN
no-reliable-ml-foundation-exists
ML lacks a reliable foundation at either the practical or theoretical level — classical and deep methods have complementary failure modes that prevent either from serving as a complete solution, while the theoretical framework that should guide choosing between them is itself undergoing fundamental revision, leaving both practical deployment and theoretical guidance in a weakened state that may require hybrid approaches. -
IN
optimization-landscape-determines-theoretical-robustness
Optimization landscape topology appears to influence how well ML theory generalizes beyond its original formulation — SVMs' convex objective guarantees global optimality and contributes to mathematical elegance, while GANs' minimax game-theoretic foundations are fragile beyond the original formulation (equilibrium equivalence breaks, Nash equilibria not guaranteed). This contrast suggests that convexity may be an important factor in theoretical robustness, though the evidence from two cases is insufficient to establish it as a necessary condition. -
IN
pan-yang-2010-canonical-transfer-learning-survey
Pan and Yang (2010) published the most-cited transfer learning survey in IEEE TKDE, defining the taxonomy of inductive, transductive, and unsupervised transfer. -
IN
practical-bridges-rest-on-dissolving-foundations
ML's practical workarounds for theoretical incompleteness are systematically built on dissolving foundations — transfer learning bridges paradigms but both endpoints rest on dissolving terrain (the classical taxonomy it formalizes is fragmenting, the modern pipelines it enables are empirically fragile and transient), while persistent manual feature engineering compensates for the manifold hypothesis's incompleteness but cannot address the reliability gap it reflects, revealing that ML's practical adaptations are parasitic on the very theoretical structures whose inadequacy they are trying to compensate for. -
IN
pragmatism-created-both-crisis-and-partial-remedy
ML's pragmatism principle is both a driver of architectural innovation and a source of theoretical fragility — the cross-field experimentation it enables contributed to discovering the manifold-geometry framework, which now provides a principled foundation for architecture design, yet this foundation does not address the reliability concerns that pragmatism's preference for empirical results over theoretical rigor helps perpetuate. -
IN
pragmatism-origin-of-permanent-reliability-gap
The permanent reliability gap originates in ML's irreducible pragmatism paradox — pragmatism enabled the discovery of mathematical necessities that validate capable ML as genuine science while simultaneously creating the crisis conditions that make reliability permanently inaccessible, meaning the very process that proved ML works is the same process that ensured it can never work safely. -
IN
pragmatism-paradox-discovers-necessities-creates-crisis
ML's pragmatism principle creates an irreducible paradox — it simultaneously enabled the discovery of deep mathematical necessities (by not requiring theoretical understanding as a precondition for adoption, allowing convergent validation) and produced the crisis dynamic (by selecting for scalability over safety), establishing that the same epistemological stance that reveals mathematical truth about ML also prevents ML from exploiting that truth for reliability. -
IN
pretrain-finetune-paradigm
The dominant transformer training paradigm is self-supervised pretraining on unlabeled data followed by supervised fine-tuning on labeled task-specific data. -
IN
pretraining-30-year-delayed-adoption
Modern self-supervised pretraining has roots in Schmidhuber's 1991 neural history compressor, which used predictive coding and self-supervised pre-training decades before the paradigm became dominant in modern deep learning — a multi-decade gap between early work and widespread adoption that suggests hardware and ecosystem readiness may play a significant role in determining when theoretical ideas achieve industrial impact. -
IN
pretraining-can-hurt-zoph-2020
Zoph et al. (2020) showed pre-training can reduce accuracy in some cases, finding self-training can outperform transfer learning when strong data augmentation is available. -
IN
pretraining-is-transfer-learning-at-scale
Modern self-supervised pretraining is transfer learning at industrial scale — the formal transfer learning framework (source domain D_S → target domain D_T) exactly describes the pretrain-then-finetune pipeline, unifying a 50-year-old theoretical concept with the dominant modern training methodology. -
IN
reliability-components-exist-but-assembly-permanently-blocked
ML possesses all the components needed for a reliable paradigm — a codified reliable methodology (SVMs), a principled design framework (manifold geometry), and complementary theoretical anchors (existence proof + architecture foundation) — but these components are permanently unassemblable because the practical bridges that could connect them rest on dissolving foundations (transfer learning spans fragmenting terrain, feature engineering compensates for incomplete theory without addressing it), creating a state where the solution exists in distributed form but no assembly pathway does. -
IN
reliability-gap-has-three-independent-impossibility-proofs
ML's reliability gap is supported by three largely independent lines of evidence operating at different levels: formal (Mitchell's definition structurally embeds the evaluation gap through proxy performance measures), economic (mathematical quality is orthogonal to evolutionary success, so reliability improvements may not survive paradigm selection), and epistemic (the crisis may be deeply intertwined with capable ML itself, suggesting reliability cannot be straightforwardly added without affecting capability) — each providing substantial independent support, collectively suggesting the gap's persistence as a structural feature rather than a solvable deficiency. -
IN
reliability-gap-permanent-not-temporary
Reliable ML appears mathematically achievable (SVMs demonstrate theory-practice unity with global optimality guarantees) yet may be systematically inaccessible — multiple avenues to reliability appear simultaneously blocked (no adequate foundation, no sufficient bridge, no effective accountability), and this blockade may not be accidental but rather deeply intertwined with capable ML itself, suggesting that the gap between what is mathematically possible and what is evolutionarily reachable could be a recurring structural feature of ML paradigms powerful enough to be useful. -
IN
reliability-knowledge-and-components-jointly-inert
ML possesses both the proven components for reliability (SVM methodology, manifold geometry, ensemble bridges — stranded across incompatible paradigms with assembly permanently blocked) AND comprehensive knowledge that these components exist and would work (mathematical necessities validated as genuine, existence proof acknowledged) — yet the components and the knowledge of how to assemble them are independently rendered inert by distinct mechanisms. -
IN
reliability-pieces-stranded-in-incompatible-paradigms
ML possesses both a codified methodology for reliable model-building (SVMs' prescriptive recipe with global optimality guarantees) and a principled framework for architecture design (manifold geometry matching data structure to inductive bias), but these assets are stranded in incompatible paradigms — SVMs' methodology cannot scale to modern problems, and manifold-based architecture design addresses geometry but not deployment reliability, meaning the field has the components of a reliable paradigm but cannot assemble them. -
IN
rnn-turing-completeness-and-svm-bayes-optimality-jointly-irrelevant
ML's two strongest mathematical results — RNNs' proven Turing-completeness (the strongest computational-theoretic result for any architecture family, made irrelevant by Transformer displacement) and SVMs' Bayes-optimal classification (the strongest statistical-theoretic result, made inaccessible by scaling barriers) — are jointly irrelevant to the field's trajectory, establishing that mathematical optimality at both the computational and statistical levels is independently orthogonal to paradigm survival. -
IN
selection-pressure-determines-reliability-over-mathematics
GANs and SVMs illustrate contrasting outcomes of different selection pressures in ML — SVMs, shaped by intellectual selection pressure, achieved strong theory-practice unity and methodological reliability, while GANs, shaped by pragmatic selection, gained unique capabilities but inherited fundamental training instability. This contrast suggests that the type of selection pressure is a primary factor in determining methodological reliability, though both cases involve sophisticated mathematics, indicating that mathematical rigor alone is insufficient without the selection environment that prioritizes it. -
IN
self-supervised-classification-debated
Whether self-supervised learning is a form of unsupervised learning or a distinct paradigm is debated among researchers -
IN
self-supervised-engine-and-symptom-of-taxonomy-dissolution
Self-supervised learning is both the engine and the symptom of paradigm taxonomy dissolution — it occupies a contested taxonomic position precisely because it is the mechanism through which modern pipelines dissolve the classical supervised/unsupervised boundary, making its own classification impossible under the framework it is dismantling. -
IN
self-supervised-learning-dominant-pretraining-paradigm
Self-supervised learning is the dominant pre-training paradigm for modern deep learning, as opposed to supervised or unsupervised learning. -
IN
self-supervised-learning-two-step-process
Self-supervised learning generates supervisory signals from data itself via augmentation/transformation, using a two-step process: a pretext task with pseudo-labels, then supervised or unsupervised fine-tuning. -
IN
self-supervised-paradigm-boundary-contested
Self-supervised learning occupies a contested taxonomic position — formally a subset of unsupervised learning, yet it has become the dominant pre-training paradigm, creating tension between its classification and its practical centrality. -
IN
self-supervised-pretraining-1991-schmidhuber
Self-supervised pre-training originated with Schmidhuber's neural history compressor in 1991, predating its use in GPT by decades -
IN
six-step-supervised-learning-procedure
The standard supervised learning procedure has six steps: determine sample type, gather training set, engineer features, select model/algorithm, train/tune with validation, evaluate on test set -
IN
supervised-learning-classification-vs-regression
Supervised learning tasks divide into classification (predicting discrete categories) and regression (predicting continuous values) -
IN
svm-and-manifold-complementary-incomplete-anchors
SVMs and the manifold hypothesis serve as complementary but individually incomplete theoretical anchors for ML — SVMs prove reliable ML is mathematically achievable (an existence proof for theory-practice unity) but cannot scale to the capability frontier, while the manifold hypothesis provides principled architecture design (geometry-matched compression) but not deployment safety — together covering theory's full scope while leaving the practical gap unbridged from either direction. -
IN
svm-artifact-of-intellectual-selection-pressure
SVMs represent what ML can achieve under intellectual rather than economic selection pressure — their unmatched theory-practice unity demonstrates the potential of mathematical rigor for producing reliable methodology, while the field's shift to economic evolutionary trajectory ensures this potential remains permanently unrealized, as economic selection systematically favors scalable pragmatism over reliable elegance. -
IN
svm-coherence-anomalous-in-pragmatic-field
SVMs' three-dimensional mathematical coherence (sparsity, equivalence, elegance) is anomalous in a field where theory is consistently violated without penalty — the most rigorous ML framework became the one that scaled least, while pragmatic architectures that violate their own mathematical prerequisites (ReLU's non-differentiability, overparameterized networks' violation of bias-variance) dominate practice. -
IN
svm-existence-proof-reliable-ml-inaccessible
SVMs suggest that reliable ML may be achievable — their unusual theory-practice unity demonstrates that mathematical rigor can produce a fully codified practical methodology — but ML's economic and research dynamics appear to select against such approaches, making reliability arguably demonstrable in principle yet difficult to reach through the field's current evolutionary trajectory. -
IN
svm-hinge-loss-bayes-optimality-deepens-existence-proof
SVMs' hinge loss recovering exactly the Bayes-optimal classifier deepens the SVM existence proof of reliable ML — reliability is not merely achievable through engineering discipline but mathematically grounded in statistical optimality theory, making the inaccessibility of this proven-optimal methodology to dominant paradigms a sharper indictment of the field's trajectory. -
IN
svm-methodology-cannot-escape-evaluation-gap
SVMs demonstrate that even ML's strongest theory-practice unity cannot escape the evaluation gap — SVMs' unmatched mathematical guarantees (convex optimization, kernel-enabled nonlinearity, codified practical methodology) exist in the training/validation domain, while evaluation itself is doubly insufficient for deployment (standard methodologies address training-test gaps but miss adversarial and bias failure modes), meaning that SVMs' mathematical guarantees, though genuine, cannot bridge the chasm between validated performance and deployment reliability. -
IN
svm-methodology-proves-reliability-achievable-but-proves-nothing-transferable
SVMs serve as evidence that reliable ML methodology may be achievable in principle (codified recipe with mathematical guarantees) while simultaneously illustrating that mathematical quality appears orthogonal to evolutionary success — together suggesting that the existence proof of reliable ML is partly self-consuming: the properties that make SVMs reliable (convexity, completeness) are among the properties associated with their inability to propagate to the paradigms that supersede them, though SVMs' guarantees themselves cannot bridge the evaluation gap between validated performance and deployment reliability. -
IN
svm-paradox-best-methodology-from-counterproductive-elegance
SVMs embody ML's deepest paradigm paradox — they are simultaneously the strongest evidence that mathematical elegance is counterproductive for paradigm survival AND the only ML framework where theoretical elegance translated into a fully codified practical methodology, suggesting that elegance's value is real but insufficient against scalability pressure. -
IN
svm-strongest-evidence-elegance-counterproductive
SVMs provide the strongest single case that mathematical elegance is actively counterproductive in ML — their anomalous three-dimensional mathematical coherence (unique in a field where theory is routinely violated without penalty) became the very property that limited their survival, as completeness created scaling barriers while pragmatic alternatives thrived precisely by lacking such constraints. -
IN
theoretical-bridges-and-practical-bridges-both-futile
ML's bridges across its reliability gap are futile at both levels — theoretical bridges (Hopfield connecting RNNs to statistical mechanics) exhibit conceptual impact exceeding practical utility while mathematical completeness is counterproductive, AND practical bridges (transfer learning, feature engineering) rest on dissolving foundations — establishing that ML cannot bridge its reliability gap from either the theoretical or practical direction. -
IN
three-concept-drift-monitoring-strategies
Three concept drift monitoring strategies are: error-based (comparing predictions to ground truth), data distribution monitoring (statistical tests on inputs), and representation monitoring (tracking hidden-layer embedding distributions). -
IN
transfer-learning-bridge-spans-dissolving-terrain
Transfer learning bridges classical and modern ML, but both sides of the bridge rest on dissolving terrain — the classical paradigm taxonomy it formalizes is dissolving, and the modern pretraining methodology it enables is empirically fragile and transient, making the bridge conceptually elegant but practically unstable. -
IN
transfer-learning-bridges-classical-and-modern-ml
Transfer learning is the conceptual bridge between classical and modern ML — it formalizes classical domain adaptation while simultaneously enabling modern LLM pipelines to dissolve paradigm boundaries, as self-supervised pretraining is precisely transfer learning operating at industrial scale across the supervised/unsupervised divide. -
IN
transfer-learning-dates-to-1976
Transfer learning dates to 1976 (Bozinovski), far earlier than the 2010s as commonly assumed. -
IN
transfer-learning-formal-definition-domain-task
Transfer learning formally involves source domain D_S with task T_S and target domain D_T with task T_T, where either the domains or tasks must differ; if both are identical it is not transfer learning. -
IN
transfer-learning-standard-for-small-data
Transfer learning (pretraining on a large dataset then fine-tuning on a small target dataset) is the standard technique when training data is limited, preventing overfitting -
IN
two-existence-proofs-of-reliability-both-inaccessible
ML possesses two independent existence proofs that reliability is achievable — SVMs prove it within ML's mathematical framework (convex optimization with Bayes-optimal recovery and global optimality guarantees) and scientific deployments prove it outside ML's framework (substituting domain-specific physical validation for absent reliability guarantees) — yet neither pathway transfers to general-purpose deployment. -
IN
unsupervised-estimates-px-supervised-pxy
Unsupervised learning estimates the a priori probability distribution p(x); supervised learning estimates the conditional distribution p(x|y) -
IN
unsupervised-learning-no-labeled-data
Unsupervised learning uses no labeled data — algorithms learn patterns exclusively from unlabeled data, in contrast to supervised learning which requires labeled datasets -
IN
unsupervised-three-algorithmic-categories
Three main algorithmic categories of unsupervised learning: clustering (k-means, DBSCAN, hierarchical), anomaly detection (local outlier factor, isolation forest), and latent variable models (EM, method of moments)