https://arxiv.org/api/ZVkaD/Amzcv7BpkRdCf2jxvjJ7YarXiv Query: search_query=&id_list=2401.00601,2401.00602,2401.00603,2401.00604,2401.00605,2401.00606,2401.00607,2401.00608,2401.00609,2401.00610,2401.00611,2401.00612,2401.00613,2401.00614,2401.00615,2401.00616,2401.00617,2401.00618,2401.00619,2401.00620,2401.00621,2401.00622,2401.00623,2401.00624,2401.00625,2401.00626,2401.00627,2401.00628,2401.00629,2401.00630,2401.00631,2401.00632,2401.00633,2401.00634,2401.00635,2401.00636,2401.00637,2401.00638,2401.00639,2401.00640,2401.00641,2401.00642,2401.00643,2401.00644,2401.00645,2401.00646,2401.00647,2401.00648,2401.00649,2401.00650,2401.00651,2401.00652,2401.00653,2401.00654,2401.00655,2401.00656,2401.00657,2401.00658,2401.00659,2401.00660,2401.00661,2401.00662,2401.00663,2401.00664,2401.00665,2401.00666,2401.00667,2401.00668,2401.00669,2401.00670,2401.00671,2401.00672,2401.00673,2401.00674,2401.00675,2401.00676,2401.00677,2401.00678,2401.00679,2401.00680,2401.00681,2401.00682,2401.00683,2401.00684,2401.00685,2401.00686,2401.00687,2401.00688,2401.00689,2401.00690,2401.00691,2401.00692,2401.00693,2401.00694,2401.00695,2401.00696,2401.00697,2401.00698,2401.00699,2401.00700&start=0&max_results=1002026-07-02T22:41:45Z1001000http://arxiv.org/abs/2401.00601v1Primal-Dual Stability in Local Optimality2023-12-31T22:56:01ZMuch is known about when a locally optimal solution depends in a single-valued Lipschitz continuous way on the problem's parameters, including tilt perturbations. Much less is known, however, about when that solution and a uniquely determined multiplier vector associated with it exhibit that dependence as a primal-dual pair. In classical nonlinear programming, such advantageous behavior is tied to the combination of the standard strong second-order sufficient condition (SSOC) for local optimality and the linear independent gradient condition (LIGC) on the active constraint gradients. But although second-order sufficient conditions have successfully been extended far beyond nonlinear programming, insights into what should replace constraint gradient independence as the extended dual counterpart have been lacking.
The exact answer is provided here for a wide range of optimization problems in finite dimensions. Behind it are advances in how coderivatives and strict graphical derivatives can be deployed. New results about strong metric regularity in solving variational inequalities and generalized equations are obtained from that as well.2023-12-31T22:56:01ZMatus BenkoR. Tyrrell Rockafellarhttp://arxiv.org/abs/2401.00622v3Federated Class-Incremental Learning with New-Class Augmented Self-Distillation2024-04-17T10:13:36ZFederated Learning (FL) enables collaborative model training among participants while guaranteeing the privacy of raw data. Mainstream FL methodologies overlook the dynamic nature of real-world data, particularly its tendency to grow in volume and diversify in classes over time. This oversight results in FL methods suffering from catastrophic forgetting, where the trained models inadvertently discard previously learned information upon assimilating new data. In response to this challenge, we propose a novel Federated Class-Incremental Learning (FCIL) method, named \underline{Fed}erated \underline{C}lass-Incremental \underline{L}earning with New-Class \underline{A}ugmented \underline{S}elf-Di\underline{S}tillation (FedCLASS). The core of FedCLASS is to enrich the class scores of historical models with new class scores predicted by current models and utilize the combined knowledge for self-distillation, enabling a more sufficient and precise knowledge transfer from historical models to current models. Theoretical analyses demonstrate that FedCLASS stands on reliable foundations, considering scores of old classes predicted by historical models as conditional probabilities in the absence of new classes, and the scores of new classes predicted by current models as the conditional probabilities of class scores derived from historical models. Empirical experiments demonstrate the superiority of FedCLASS over four baseline algorithms in reducing average forgetting rate and boosting global accuracy.2024-01-01T00:54:02Z9 pages, 2 figures, 4 tablesZhiyuan WuTianliu HeSheng SunYuwei WangMin LiuBo GaoXuefeng Jianghttp://arxiv.org/abs/2401.00640v3On the Law of Large Numbers and Convergence Rates for the Discrete Fourier Transform of Random Fields2024-01-19T21:26:58ZWe study the Marcinkiewicz-Zygmund strong law of large numbers for the cubic partial sums of the discrete Fourier transform of random fields. We establish Marcinkiewicz-Zygmund types rate of convergence for the discrete Fourier transform of random fields under weaker conditions than identical distribution.2024-01-01T02:37:21Z Vishakhahttp://arxiv.org/abs/2401.00646v1High magnetic field phase diagram and weak FM breaking in (Ni0.93Co0.07)3V2O82024-01-01T03:20:27ZWe present magnetostriction and thermal expansion measurements on multiferroic (Ni0.93Co0.07)3V2O8. The high field phase diagrams up to 33 T along the a, b and c directions are built. For H//a, as the magnetic field increases, two intermediate phases appear between the incommensurate phase and the paramagnetic phase at about 7 K, and then a magnetically induced phase appears above the paramagnetic phase. For H//b,thermal expansion measurement indicates a mutation in the spin lattice coupling of the high field phases. The interlaced phase boundary suggests a mixed state in the optical high field phase. For H//c, an intermediate phase between the commensurate phase and the incommensurate phase is detected. A nonlinear boundary between the intermediate phase and the low temperature incommensurate phase, and a clear boundary between the commensurate phase and the paramagnetic phase are found. These results indicate that doping Co2+ breaks the weak ferromagnetic moment of the commensurate phase, which exists in the parent compound Ni3V2O8 and (Ni0.9Co0.1)3V2O8. This nonlinear influence reflects complicated spin modulation in Ni3V2O8 by doping Co2+.2024-01-01T03:20:27Z7 pages, 4 figuresPhys. Rev. B 108, 214108(2023)Jiating WuMinjie ZhangKe ShiHuxin YinYuyan HanLansheng LingWei TongChuanying XiLi PiZhaosheng Wang10.1103/PhysRevB.108.214108http://arxiv.org/abs/2401.00650v1Automated Invariant Generation for Solidity Smart Contracts2024-01-01T03:37:30ZSmart contracts are computer programs running on blockchains to automate the transaction execution between users. The absence of contract specifications poses a real challenge to the correctness verification of smart contracts. Program invariants are properties that are always preserved throughout the execution, which characterize an important aspect of the program behaviors. In this paper, we propose a novel invariant generation framework, INVCON+, for Solidity smart contracts. INVCON+ extends the existing invariant detector, InvCon, to automatically produce verified contract invariants based on both dynamic inference and static verification. Unlike INVCON+, InvCon only produces likely invariants, which have a high probability to hold, yet are still not verified against the contract code. Particularly, INVCON+ is able to infer more expressive invariants that capture richer semantic relations of contract code. We evaluate INVCON+ on 361 ERC20 and 10 ERC721 real-world contracts, as well as common ERC20 vulnerability benchmarks. The experimental results indicate that INVCON+ efficiently produces high-quality invariant specifications, which can be used to secure smart contracts from common vulnerabilities.2024-01-01T03:37:30ZYe LiuNanyang Technological University, SingaporeChengxuan ZhangNanyang Technological University, SingaporeYi Li.Nanyang Technological University, Singaporehttp://arxiv.org/abs/2401.00671v1Large deviation principle for a two-time-scale McKean-Vlasov model with jumps2024-01-01T05:28:12ZThis work focus on the large deviation principle for a two-time scale McKean-Vlasov system with jumps. Based on the variational framework of the McKean-Vlasov system with jumps, it is turned into weak convergence for the controlled system. Unlike general two-time scale system, the controlled McKean-Vlasov system is related to the law of the original system, which causes difficulties in qualitative analysis. In solving this problem, employing asymptotics of the original system and a Khasminskii-type averaging principle together is efficient. Finally, it is shown that the limit is related to the Dirac measure of the solution to the ordinary differential equation.2024-01-01T05:28:12ZXiaoyu YangYong Xuhttp://arxiv.org/abs/2401.00690v1Benchmarking Large Language Models on Controllable Generation under Diversified Instructions2024-01-01T07:35:31ZWhile large language models (LLMs) have exhibited impressive instruction-following capabilities, it is still unclear whether and to what extent they can respond to explicit constraints that might be entailed in various instructions. As a significant aspect of LLM alignment, it is thus important to formulate such a specialized set of instructions as well as investigate the resulting behavior of LLMs. To address this vacancy, we propose a new benchmark CoDI-Eval to systematically and comprehensively evaluate LLMs' responses to instructions with various constraints. We construct a large collection of constraints-attributed instructions as a test suite focused on both generalization and coverage. Specifically, we advocate an instruction diversification process to synthesize diverse forms of constraint expression and also deliberate the candidate task taxonomy with even finer-grained sub-categories. Finally, we automate the entire evaluation process to facilitate further developments. Different from existing studies on controllable text generation, CoDI-Eval extends the scope to the prevalent instruction-following paradigm for the first time. We provide extensive evaluations of representative LLMs (e.g., ChatGPT, Vicuna) on CoDI-Eval, revealing their limitations in following instructions with specific constraints and there is still a significant gap between open-source and commercial closed-source LLMs. We believe this benchmark will facilitate research into improving the controllability of LLMs' responses to instructions. Our data and code are available at https://github.com/Xt-cyh/CoDI-Eval.2024-01-01T07:35:31ZAccepted to AAAI 2024Yihan ChenBenfeng XuQuan WangYi LiuZhendong Maohttp://arxiv.org/abs/2401.00697v13D non-LTE abundance analyses of late-type stars2024-01-01T08:25:33ZThe chemical compositions of stars encode the history of the universe and are thus fundamental for advancing our knowledge of astrophysics and cosmology. However, measurements of elemental abundances ratios, and our interpretations of them, strongly depend on the physical assumptions that dictate the generation of synthetic stellar spectra. Three-dimensional radiation-hydrodynamic (3D RHD) ``box-in-a-star'' simulations of stellar atmospheres offer a more realistic representation of surface convection occurring in late-type stars compared to traditional one-dimensional (1D) hydrostatic models. As evident from a multitude of observational tests, the coupling of 3D RHD models with line-formation in non-local thermodynamic equilibrium (non-LTE) today provides a solid foundation for abundance analysis for many elements. This review describes the ongoing and transformational work to advance the state-of-the-art and replace 1D LTE spectrum synthesis with its 3D non-LTE counterpart. In summary:
1) 3D and non-LTE effects are intricately coupled and consistent modelling thereof is necessary for high-precision abundances, which is currently feasible for individual elements in large surveys. Mean 3D (<3D>) models are not adequate as substitutes.
2) The solar abundance debate is presently dominated by choices and systematic uncertainties that are not specific to 3D non-LTE modelling.
3) 3D non-LTE abundance corrections have a profound impact on our understanding of FGK-type stars, exoplanets, and the nucleosynthetic origins of the elements.2024-01-01T08:25:33ZTo appear in Annual Reviews of Astronomy and Astrophysics (65 pages, 13 figures)Karin LindAnish Mayur Amarsihttp://arxiv.org/abs/2401.00605v2Distributed Multi-Object Tracking Under Limited Field of View Heterogeneous Sensors with Density Clustering2024-09-11T01:22:51ZWe consider the problem of tracking multiple, unknown, and time-varying numbers of objects using a distributed network of heterogeneous sensors. In an effort to derive a formulation for practical settings, we consider limited and unknown sensor field-of-views (FoVs), sensors with limited local computational resources and communication channel capacity. The resulting distributed multi-object tracking algorithm involves solving an NP-hard multidimensional assignment problem either optimally for small-size problems or sub-optimally for general practical problems. For general problems, we propose an efficient distributed multi-object tracking algorithm that performs track-to-track fusion using a clustering-based analysis of the state space transformed into a density space to mitigate the complexity of the assignment problem. The proposed algorithm can more efficiently group local track estimates for fusion than existing approaches. To ensure we achieve globally consistent identities for tracks across a network of nodes as objects move between FoVs, we develop a graph-based algorithm to achieve label consensus and minimise track segmentation. Numerical experiments with synthetic and real-world trajectory datasets demonstrate that our proposed method is significantly more computationally efficient than state-of-the-art solutions, achieving similar tracking accuracy and bandwidth requirements but with improved label consistency.2023-12-31T23:12:44ZFor data and code see, https://github.com/AdelaideAuto-IDLab/Distributed-limitedFoV-CDPMOTSignal Processing (2024)Fei ChenHoa Van NguyenAlex S. LeongSabita PanickerRobin BakerDamith C. Ranasinghehttp://arxiv.org/abs/2401.00654v3Global rigidity for some partially hyperbolic abelian actions with 1-dimensional center2024-10-04T02:26:24ZWe obtain a global rigidity result for abelian partially hyperbolic higher rank actions on certain $2-$step nilmanifolds $X_Γ$. We show that, under certain natural assumptions, all such actions are $C^{\infty}-$conjugated to an affine model. As a consequence, we obtain a centralizer rigidity result, classifying all possible centralizers for any $C^{1}-$small perturbation of an irreducible, affine partially hyperbolic map on $X_Γ$. Along the way, we also prove two results of independent interest. We describe fibered partially hyperbolic diffeomorphisms on $X_Γ$ and we show that topological conjugacies between partially hyperbolic actions and higher rank affine actions are $C^{\infty}$.2024-01-01T03:49:36Z54 pages, 1 figure. Changes in introduction and formatSven Sandfeldthttp://arxiv.org/abs/2401.00674v1Biphenylene Nanoribbon as Promising Electrocatalyst for Hydrogen Evolution2024-01-01T05:44:00ZDesigning efficient, metal free, and in-expensive catalyst for electrochemical hydrogen evolution reaction (HER) is crucial for large scale clean and green energy production. Recently synthesized 1D Biphenylene nanoribbons (BPRs) display few exciting properties originating from the unique co-existence of 4, 6 and 8 coordinated carbon rings. Here, we present a first principles calculation of the electronic structure and electrocatalytic activity of various sized N-BPRs (N indicates the width). The electrocatalytic performance is evaluated based on several descriptors including electronic property, carrier mobility, Gibbs free energy ($Δ$G$_{HER}$), exchange current density etc. The electronic properties are crucially sensitive to the width (N) of BPRs transiting it from semiconducting to metallic nature at N=18. The p$_z$ orbitals of C-atoms from the central tetragonal rings are mainly responsible for the decrease in band gap with increasing width. The room temperature electron mobility is found to be as high as $\sim6.3\times$10$^4$ cm$^2$V$^{-1}$s$^{-1}$, while hole mobility is relatively lower in magnitude. Gibbs free energy also depends sensitively on the width of BPRs as evident from the p-band center analysis. We propose 15-BPR as the most promising candidate for electrocatalytic activity with extremely small overpotential ($\approx$0.005 V) and a high exchange current density, much better than the state-of-the-art Pt(111). A close inspection of the elementary reactions of HER (Volmer, Heyrovsky and Tafel) confirms Volmer-Tafel mechanism to be most dominant on 15-BPR with Tafel as the rate-determining step with a barrier of 0.56 eV. The present study provides a deeper insight into the excellent HER catalytic activity of a newly synthesized BPR which is inexpensive and is expected to hasten experimental validation towards H$_2$ production.2024-01-01T05:44:00Z10 pages, 5 figures, 2 Tables, and Supporting Information (total 8-pages with 7 figures)Radha N SomaiyaZicong Marvin WongBrahmananda ChakrabortyTeck Leong TanAftab Alam10.1103/PhysRevApplied.21.054007http://arxiv.org/abs/2401.00686v2Transit cosmological models in Myrzakulov F(R,T) gravity theory2024-02-19T06:29:57ZIn the present paper, we investigate some exact cosmological models in Myrzakulov $F(R,T)$ gravity theory. We have considered the arbitrary function $F(R, T)=R+λT$ where $λ$ is an arbitrary constant, $R, T$ are respectively, the Ricci-scalar curvature and the torsion. We have solved the field equations in a flat FLRW spacetime manifold for Hubble parameter and using the MCMC analysis, we have estimated the best fit values of model parameters with $1-σ, 2-σ, 3-σ$ regions, for two observational datasets like $H(z)$ and Pantheon SNe Ia datasets. Using these best fit values of model parameters, we have done the result analysis and discussion of the model. We have found a transit phase decelerating-accelerating universe model with transition redshifts $z_{t}=0.4438_{-0.790}^{+0.1008}, 0.3651_{-0.0904}^{+0.1644}$. The effective dark energy equation of state varies as $-1\leω_{de}\le-0.5176$ and the present age of the universe is found as $t_{0}=13.8486_{-0.0640}^{+0.1005}, 12.0135_{-0.2743}^{+0.6206}$ Gyrs, respectively for two datasets.2024-01-01T07:20:42Z22 pages, 11 figuresEur. Phys. J. C 84, 534 (2024)Dinesh Chandra MauryaRatbay Myrzakulov10.1140/epjc/s10052-024-12904-5http://arxiv.org/abs/2401.00667v2Channelling Multimodality Through a Unimodalizing Transport: Warp-U Sampler and Stochastic Bridge Sampling Estimator2025-03-08T02:53:21ZMonte Carlo integration is a powerful tool for scientific and statistical computation, but faces significant challenges when the integrand is a multi-modal distribution, even when the mode locations are known. This work introduces novel Monte Carlo sampling and integration estimation strategies for the multi-modal context by leveraging a generalized version of the stochastic Warp-U transformation Wang et al. [2022]. We propose two flexible classes of Warp-U transformations, one based on a general location-scale-skew mixture model and a second using neural ordinary differential equations. We develop an efficient sampling strategy called Warp-U sampling, which applies a Warp-U transformation to map a multi-modal density into a uni-modal one, then inverts the transformation with injected stochasticity. In high dimensions, our approach relies on information about the mode locations, but requires minimal tuning and demonstrates better mixing properties than conventional methods with identical mode information. To improve normalizing constant estimation once samples are obtained, we propose a stochastic Warp-U bridge sampling estimator, which we demonstrate has higher asymptotic precision per CPU second compared to the original approach proposed by Wang et al. [2022]. We also establish the ergodicity of our sampling algorithm. The effectiveness and current limitations of our methods are illustrated through simulation studies and an application to exoplanet detection.2024-01-01T05:10:06ZFei DingShiyuan HeDavid E. JonesXiao-Li Menghttp://arxiv.org/abs/2401.00691v5Stochastic Gradient Descent for Nonparametric Additive Regression2025-12-30T17:23:40ZThis paper introduces an iterative algorithm for training nonparametric additive models that enjoys favorable memory storage and computational requirements. The algorithm can be viewed as the functional counterpart of stochastic gradient descent, applied to the coefficients of a truncated basis expansion of the component functions. We show that the resulting estimator satisfies an oracle inequality that allows for model mis-specification. In the well-specified setting, by choosing the learning rate carefully across three distinct stages of training, we demonstrate that its risk is minimax optimal in terms of the dependence on both the dimensionality of the data and the size of the training sample. Unlike past work, we also provide polynomial convergence rates even when the covariates do not have full support on their domain.2024-01-01T08:03:52ZXin ChenJason M. Klusowskihttp://arxiv.org/abs/2401.00615v2Test ideals in mixed characteristic: a unified theory up to perturbation2025-07-08T23:20:09ZLet $X$ be an integral scheme of finite type over a complete DVR of mixed characteristic. We provide a definition of a test ideal which agrees with the multiplier ideal after inverting $p$, is computed from a sufficiently large alteration, agrees with previous mixed characteristic BCM test ideals after completing at any point of residue characteristic $p$ (up to small perturbation), and which satisfies the full suite of expected properties of a multiplier or test ideal. This object is obtained via the $p$-adic Riemann-Hilbert functor.2024-01-01T00:05:48Z117 pages, major revision, improved results (removed completeness hypothesis from base DVR, better perturbation/test elements), added appendix by Rankeya Datta, other minor changes and typos corrected, comments welcomeBhargav BhattLinquan MaZsolt PatakfalviKarl SchwedeKevin TuckerJoe WaldronJakub WitaszekRankeya Dattahttp://arxiv.org/abs/2401.00620v1A Borel-Pompeiu formula in a $(q,q')$-model of quaternionic analysis2024-01-01T00:20:52ZThe study of $ψ-$hyperholomorphic functions defined on domains in $\mathbb R^4$ with values in $\mathbb H$, namely null-solutions of the $ψ-$Fueter operator, is a topic which captured great interest in quaternionic analysis. This class of functions is more general than that of Fueter regular functions. In the setting of $(q,q')-$calculus, also known as post quantum calculus, we introduce a deformation of the $ψ-$Fueter operator written in terms of suitable difference operators, which reduces to a deformed $q$ calculus when $q'=1$. We also prove the Stokes and Borel-Pompeiu formulas in this context. This work is the first investigation of results in quaternionic analysis in the setting of the $(q,q')-$calculus theory.2024-01-01T00:20:52ZJosé Oscar González-CervantesJuan Bory-ReyesIrene Sabadinihttp://arxiv.org/abs/2401.00623v1Normalized solutions to the Chern-Simons-Schrödinger system: the supercritical case2024-01-01T00:56:44ZWe are concerned with the existence of normalized solutions for a class of generalized Chern-Simons-Schrödinger type problems with supercritical exponential growth
$$ -Δu +λu+A_0 u+\sum\limits_{j=1}^2A_j^2 u=f(u),\quad
\partial_1A_2-\partial_2A_1=-\frac{1}{2}|u|^2,\quad \partial_1A_1+\partial_2A_2=0,\quad
\partial_1A_0=A_2|u|^2,\quad \partial_2A_0=-A_1|u|^2,\quad
\int_{\mathbb{R}^2}|u|^2dx=a^2, $$ where $a\neq0$, $λ\in \mathbb{R}$ is known as the Lagrange multiplier and $f\in C^1(\mathbb{R})$ denotes the nonlinearity that fulfills the supercritical exponential growth in the Trudinger-Moser sense at infinity. Under suitable assumptions, combining the constrained minimization approach together with the homotopy stable family and elliptic regularity theory, we obtain that the problem has at least a ground state solution.2024-01-01T00:56:44Z39 pagesLiejun ShenMarco Squassinahttp://arxiv.org/abs/2401.00637v1Nonlinear vibration of a dipteran flight robot system with rotational geometric nonlinearity2024-01-01T02:25:39ZThe dipteran flight mechanism of the insects is commonly used to design the nonlinear flight robot system. However, the dynamic response of the click mechanism of the nonlinear robot system with multiple stability still unclear. In this paper, a novel dipteran robot model with click mechanism proposed based on the multiple stability of snap-through buckling. The motion of equation of the nonlinear flight robot system is obtained by using the Euler-Lagrange equation. The nonlinear potential energy, the elastic force, equilibrium bifurcation, as well as equilibrium stability are investigated to show the multiple stability characteristics. The transient sets of bifurcation and persistent set of regions in the system parameter plane and the corresponding phase portraits are obtained with multiple stability of single and double well behaviors. Then, the periodic free vibration response are defined by the analytical solution of three kinds of elliptical functions, as well as the amplitude frequency responses are investigated by numerical integration. Based on the topological equivalent method, the chaotic thresholds of the homo-clinic orbits for the chaotic vibration of harmonic forced robot system are derived to show the chaotic parametric condition. Finally, the prototype of nonlinear flapping robot is manufactured and the experimental system is setup. The nonlinear static moment of force curves, periodic response and dynamic flight vibration of dipteran robot system are carried out. It is shown that the test results are agree well with the theoretical analysis and numerical simulation. Those result have the potential application for the structure design of the efficient flight robot.2024-01-01T02:25:39Z30 pages, 24 figureYanwei HanZijian Zhanghttp://arxiv.org/abs/2401.00641v3Hierarchical Bayesian Modeling for Time-Dependent Inverse Uncertainty Quantification2024-03-26T01:16:28ZThis paper introduces a novel hierarchical Bayesian model specifically designed to address challenges in Inverse Uncertainty Quantification (IUQ) for time-dependent problems in nuclear Thermal Hydraulics (TH) systems. The unique characteristics of time-dependent data, such as high dimensionality and correlation in model outputs requires special attention in the IUQ process. By integrating Gaussian Processes (GP) with Principal Component Analysis (PCA), we efficiently construct surrogate models that effectively handle the complexity of dynamic TH systems. Additionally, we incorporate Neural Network (NN) models for time series regression, enhancing the computational accuracy and facilitating derivative calculations for efficient posterior sampling using the Hamiltonian Monte Carlo Method - No U-Turn Sampler (NUTS).
We demonstrate the effectiveness of this hierarchical Bayesian approach using the transient experiments in the PSBT benchmark. Our results show improved estimates of Physical Model Parameters' posterior distributions and a reduced tendency for over-fitting, compared to conventional single-level Bayesian models. This approach offers a promising framework for extending IUQ to more complex, time-dependent problems.2024-01-01T02:39:54ZChen Wanghttp://arxiv.org/abs/2401.00642v1Predicting Anti-microbial Resistance using Large Language Models2024-01-01T03:04:14ZDuring times of increasing antibiotic resistance and the spread of infectious diseases like COVID-19, it is important to classify genes related to antibiotic resistance. As natural language processing has advanced with transformer-based language models, many language models that learn characteristics of nucleotide sequences have also emerged. These models show good performance in classifying various features of nucleotide sequences. When classifying nucleotide sequences, not only the sequence itself, but also various background knowledge is utilized. In this study, we use not only a nucleotide sequence-based language model but also a text language model based on PubMed articles to reflect more biological background knowledge in the model. We propose a method to fine-tune the nucleotide sequence language model and the text language model based on various databases of antibiotic resistance genes. We also propose an LLM-based augmentation technique to supplement the data and an ensemble method to effectively combine the two models. We also propose a benchmark for evaluating the model. Our method achieved better performance than the nucleotide sequence language model in the drug resistance class prediction.2024-01-01T03:04:14ZHyunwoo YooBahrad SokhansanjJames R. BrownGail Rosenhttp://arxiv.org/abs/2401.00647v1Tuning Thermal Conductivity of Hybrid Perovskites through Halide Alloying2024-01-01T03:20:45ZTuning the thermal transport properties of hybrid halide perovskites is critical for their applications in optoelectronics, thermoelectrics, and photovoltaics. Here, we demonstrate an effective strategy to modulate the thermal transport property of hybrid perovskites by halide alloying. A highly tunable thermal conductivity of mixed-halide hybrid perovskites is achieved due to halide-alloying and structural distortion. Our experimental measurements show that the room temperature thermal conductivity of MAPb(BrxI1-x)3 (x = 0-1) can be largely modulated from 0.27 W/mK (x = 0.5) to 0.47 W/mK (x = 1). Molecular dynamics simulations further demonstrate that the thermal conductivity reduction of hybrid halide perovskites results from the suppression of the mean free paths of the low-frequency acoustic and optical phonons. It is found that halide alloying and the induced structural distortion can largely increase the scatterings of optical and acoustic phonons, respectively. The confined diffusion of MA+ cations in the octahedra cage is found to act as an additional thermal transport channel in hybrid perovskites and can contribute around 10-20% of the total thermal conductivity. Our findings provide a strategy for tailoring the thermal transport in hybrid halide perovskites which may largely benefit their related applications.2024-01-01T03:20:45ZGuang WangHongzhao FanZhongwei ChenYufei GaoZuankai WangZhigang LiHaipeng LuYanguang Zhouhttp://arxiv.org/abs/2401.00661v1Personalized Dynamic Pricing Policy for Electric Vehicles: Reinforcement learning approach2024-01-01T04:20:30ZWith the increasing number of fast-electric vehicle charging stations (fast-EVCSs) and the popularization of information technology, electricity price competition between fast-EVCSs is highly expected, in which the utilization of public and/or privacy-preserved information will play a crucial role. Self-interest electric vehicle (EV) users, on the other hand, try to select a fast-EVCS for charging in a way to maximize their utilities based on electricity price, estimated waiting time, and their state of charge. While existing studies have largely focused on finding equilibrium prices, this study proposes a personalized dynamic pricing policy (PeDP) for a fast-EVCS to maximize revenue using a reinforcement learning (RL) approach. We first propose a multiple fast-EVCSs competing simulation environment to model the selfish behavior of EV users using a game-based charging station selection model with a monetary utility function. In the environment, we propose a Q-learning-based PeDP to maximize fast-EVCS' revenue. Through numerical simulations based on the environment: (1) we identify the importance of waiting time in the EV charging market by comparing the classic Bertrand competition model with the proposed PeDP for fast-EVCSs (from the system perspective); (2) we evaluate the performance of the proposed PeDP and analyze the effects of the information on the policy (from the service provider perspective); and (3) it can be seen that privacy-preserved information sharing can be misused by artificial intelligence-based PeDP in a certain situation in the EV charging market (from the customer perspective).2024-01-01T04:20:30ZSangjun BaeBalazs KulcsarSebastien Groshttp://arxiv.org/abs/2401.00678v1General-purpose foundation models for increased autonomy in robot-assisted surgery2024-01-01T06:15:16ZThe dominant paradigm for end-to-end robot learning focuses on optimizing task-specific objectives that solve a single robotic problem such as picking up an object or reaching a target position. However, recent work on high-capacity models in robotics has shown promise toward being trained on large collections of diverse and task-agnostic datasets of video demonstrations. These models have shown impressive levels of generalization to unseen circumstances, especially as the amount of data and the model complexity scale. Surgical robot systems that learn from data have struggled to advance as quickly as other fields of robot learning for a few reasons: (1) there is a lack of existing large-scale open-source data to train models, (2) it is challenging to model the soft-body deformations that these robots work with during surgery because simulation cannot match the physical and visual complexity of biological tissue, and (3) surgical robots risk harming patients when tested in clinical trials and require more extensive safety measures. This perspective article aims to provide a path toward increasing robot autonomy in robot-assisted surgery through the development of a multi-modal, multi-task, vision-language-action model for surgical robots. Ultimately, we argue that surgical robots are uniquely positioned to benefit from general-purpose models and provide three guiding actions toward increased autonomy in robot-assisted surgery.2024-01-01T06:15:16ZSamuel SchmidgallJi Woong KimAlan KuntzAhmed Ezzat GhaziAxel Kriegerhttp://arxiv.org/abs/2401.00693v2Decoupling Limits in Effective Field Theories via Higher Dimensional Operators2024-02-11T21:05:25ZNon-decoupling effects of heavy scalars and vector fields play an important role in the indirect search of Beyond the Standard Model (BSM) physics at the LHC. By exploiting some new differential equations for the 1-PI amplitudes, we show that such non-decoupling effects are absent for quite a general class of effective field theories involving dimension six two-derivatives and dimension eight four-derivatives operators, once resummation in certain BSM couplings is taken into account and some particular regimes of the relevant couplings are considered.2024-01-01T08:15:37Z22 pages. Final version published in the journalUniverse 2024, 10(2), 85Andrea QuadriINFN, Milan10.3390/universe10020085http://arxiv.org/abs/2401.00694v1A note on a question of Shioda about integral sections2024-01-01T08:18:55ZWe consider a rational elliptic surface with a relatively minimal fibration. We compute the number of integral sections in the above rational elliptic surface. As an application, we obtain an estimate of polynomial solutions of some equations.2024-01-01T08:18:55Z23 pages, 2 figuresJia-Li Mohttp://arxiv.org/abs/2401.00602v1An analysis of protesting activity and trauma through mathematical and statistical models2023-12-31T22:58:49ZThe effect that different police protest management methods have on protesters' physical and mental trauma is still not well understood and is a matter of debate. In this paper, we take a two-pronged approach to gain insight into this issue. First, we perform statistical analysis on time series data of protests provided by ACLED and spanning the period of time from January 1, 2020, until March 13, 2021. We observe that the use of kinetic impact projectiles is associated with more protests in subsequent days and is also a better predictor of the number of deaths in subsequent deaths than the number of protests, concluding that the use of non-lethal weapons seems to have an inflammatory rather than suppressive effect on protests. Next, we provide a mathematical framework to model modern, but well-established psychological and sociological research on compliance theory and crowd dynamics. Our results show that understanding the heterogeneity of the crowd is key for protests that lead to a reduction of social tension and minimization of physical and mental trauma in protesters.2023-12-31T22:58:49ZThis paper was published in 2023Crime Science, 12(1):17, 2023Nancy RodriguezDavid White10.1186/s40163-023-00197-0http://arxiv.org/abs/2401.00603v1Intraday Trading Algorithm for Predicting Cryptocurrency Price Movements Using Twitter Big Data Analysis2023-12-31T23:03:16ZCryptocurrencies have emerged as a novel financial asset garnering significant attention in recent years. A defining characteristic of these digital currencies is their pronounced short-term market volatility, primarily influenced by widespread sentiment polarization, particularly on social media platforms such as Twitter. Recent research has underscored the correlation between sentiment expressed in various networks and the price dynamics of cryptocurrencies. This study delves into the 15-minute impact of informative tweets disseminated through foundation channels on trader behavior, with a focus on potential outcomes related to sentiment polarization. The primary objective is to identify factors that can predict positive price movements and potentially be leveraged through a trading algorithm. To accomplish this objective, we conduct a conditional examination of return and excess return rates within the 15 minutes following tweet publication. The empirical findings reveal statistically significant increases in return rates, particularly within the initial three minutes following tweet publication. Notably, adverse effects resulting from the messages were not observed. Surprisingly, sentiments were found to have no discerni-ble impact on cryptocurrency price movements. Our analysis further identifies that inves-tors are primarily influenced by the quality of tweet content, as reflected in the choice of words and tweet volume. While the basic trading algorithm presented in this study does yield some benefits within the 15-minute timeframe, these benefits are not statistically significant. Nevertheless, it serves as a foundational framework for potential enhance-ments and further investigations.2023-12-31T23:03:16ZVahidin JeleskovicStephen Mackayhttp://arxiv.org/abs/2401.00610v1A High School Camp on Algorithms and Coding in Jamaica2023-12-31T23:52:03ZThis is a report on JamCoders, a four-week long computer-science camp for high school students in Jamaica. The camp teaches college-level coding and algorithms, and targets academically excellent students in grades 9--11 (ages 14--17). Qualitative assessment shows that the camp was, in general terms, a success. We reflect on the background and academic structure of the camp and share key takeaways on designing and operating a successful camp. We analyze data collected before, during and after the camp and map the effects of demographic differences on student performance in camp. We conclude with a discussion on possible improvements on our approach.2023-12-31T23:52:03ZTo appear in Proceedings of the 55th ACM Technical Symposium on Computer Science Education (SIGCSE), 2024Daniel T. FokumZaria Chen ShuiKerene WrightOrr ParadiseGunjan MansinghDaniel Coorehttp://arxiv.org/abs/2401.00612v3On the relations between Auerbach or almost Auberbach Markushevich systems and Schauder bases2024-06-10T02:42:43ZWe establish that the summability of the series $\sum\varepsilon_n$ is the necessary and sufficient criterion ensuring that every $(1+\varepsilon_n)$ Markushevich basis in a separable Hilbert space is a Riesz basis. Further we show that if $n\varepsilon_n\to \infty$, then in $\ell_2$ there exists a $(1+\varepsilon_n)$ Markushevich basis that under any permutation is non-equivalent to a Schauder basis. We extend this result to any separable Banach space. Finally we provide examples of Auerbach bases in 1-symmetric separable Banach spaces whose no permutations are equivalent to any Schauder basis or (depending on the space) any unconditional Schauder basis.2023-12-31T23:57:44ZWe improved the presentation of the proof of Theorem 1.3 and corrected some typos. There are no essential changes except the removal of the last sentence of Example 8.5 which was not correct. The main statement of Example 8.5 remains unchanged. We slightly reformulated the AcknowledgmentsBeata RandrianantoaninaMichał WojciechowskiPavel Zatitskiihttp://arxiv.org/abs/2401.00687v2Constraining MeV to 10 GeV majoron by Big Bang Nucleosynthesis2024-07-01T04:32:51ZWe estimate the Big Bang nucleosynthesis (BBN) constraint on the majoron in the mass range between $1\,{\rm MeV}$ to $10\,{\rm GeV}$ which dominantly decays into the standard model neutrinos. When the majoron lifetime is shorter than $1\,{\rm sec}$, the injected neutrinos mainly heat up background plasma, which alters the relation between photon temperature and background neutrino temperature. For a lifetime longer than $1\,{\rm sec}$, most of the injected neutrinos directly contribute to the protons-to-neutrons conversion. In both cases, deuterium and helium abundances are enhanced, while the constraint from the deuterium is stronger than that from the helium. $^7{\rm Li}$ abundance gets decreased as a consequence of additional neutrons, but the parameter range that fits the observed $^7{\rm Li}$ abundance is excluded by the deuterium constraint. We also estimate other cosmological constraints and compare them with the BBN bound.2024-01-01T07:23:54Z12 pages, 3 figures, v2: major revision made by including correct n/p conversion induced by modified thermal neutrinos, slight changes in the result, references added, version accepted in PRDSanghyeon ChangSougata GangulyTae Hyun JungTae-Sun ParkChang Sub Shinhttp://arxiv.org/abs/2401.00629v2Adversarially Trained Weighted Actor-Critic for Safe Offline Reinforcement Learning2024-10-31T07:06:05ZWe propose WSAC (Weighted Safe Actor-Critic), a novel algorithm for Safe Offline Reinforcement Learning (RL) under functional approximation, which can robustly optimize policies to improve upon an arbitrary reference policy with limited data coverage. WSAC is designed as a two-player Stackelberg game to optimize a refined objective function. The actor optimizes the policy against two adversarially trained value critics with small importance-weighted Bellman errors, which focus on scenarios where the actor's performance is inferior to the reference policy. In theory, we demonstrate that when the actor employs a no-regret optimization oracle, WSAC achieves a number of guarantees: (i) For the first time in the safe offline RL setting, we establish that WSAC can produce a policy that outperforms any reference policy while maintaining the same level of safety, which is critical to designing a safe algorithm for offline RL. (ii) WSAC achieves the optimal statistical convergence rate of $1/\sqrt{N}$ to the reference policy, where $N$ is the size of the offline dataset. (iii) We theoretically show that WSAC guarantees a safe policy improvement across a broad range of hyperparameters that control the degree of pessimism, indicating its practical robustness. Additionally, we offer a practical version of WSAC and compare it with existing state-of-the-art safe offline RL algorithms in several continuous control environments. WSAC outperforms all baselines across a range of tasks, supporting the theoretical results.2024-01-01T01:44:58ZHonghao WeiXiyue PengArnob GhoshXin Liuhttp://arxiv.org/abs/2401.00692v3Self-supervised learning for skin cancer diagnosis with limited training data2024-11-26T05:29:17ZEarly cancer detection is crucial for prognosis, but many cancer types lack large labelled datasets required for developing deep learning models. This paper investigates self-supervised learning (SSL) as an alternative to the standard supervised pre-training on ImageNet for scenarios with limited training data using a deep learning model (ResNet-50). We first demonstrate that SSL pre-training on ImageNet (via the Barlow Twins SSL algorithm) outperforms supervised pre-training (SL) using a skin lesion dataset with limited training samples. We then consider \textit{further} SSL pre-training (of the two ImageNet pre-trained models) on task-specific datasets, where our implementation is motivated by supervised transfer learning. This approach significantly enhances initially SL pre-trained models, closing the performance gap with initially SSL pre-trained ones. Surprisingly, further pre-training on just the limited fine-tuning data achieves this performance equivalence. Linear probe experiments reveal that improvement stems from enhanced feature extraction. Hence, we find that minimal further SSL pre-training on task-specific data can be as effective as large-scale SSL pre-training on ImageNet for medical image classification tasks with limited labelled data. We validate these results on an oral cancer histopathology dataset, suggesting broader applicability across medical imaging domains facing labelled data scarcity.2024-01-01T08:11:38ZHamish HaggertyRohitash Chandrahttp://arxiv.org/abs/2401.00688v2Inference and Visualization of Community Structure in Attributed Hypergraphs Using Mixed-Membership Stochastic Block Models2025-05-04T16:33:59ZHypergraphs represent complex systems involving interactions among more than two entities and allow the investigation of higher-order structure and dynamics in complex systems. Node attribute data, which often accompanies network data, can enhance the inference of community structure in complex systems. While mixed-membership stochastic block models have been employed to infer community structure in hypergraphs, they complicate the visualization and interpretation of inferred community structure by assuming that nodes may possess soft community memberships. In this study, we propose a framework, HyperNEO, that combines mixed-membership stochastic block models for hypergraphs with dimensionality reduction methods. Our approach generates a node layout that largely preserves the community memberships of nodes. We evaluate our framework on both synthetic and empirical hypergraphs with node attributes. We expect our framework will broaden the investigation and understanding of higher-order community structure in complex systems.2024-01-01T07:31:32ZSocial Network Analysis and Mining. Vol. 15, Article No. 5 (2025)Kazuki NakajimaTakeaki Uno10.1007/s13278-025-01440-zhttp://arxiv.org/abs/2401.00645v2Density bounds for unit ball packings relative to their outer parallel domains2024-04-19T14:44:50ZWe prove that the highest density of non-overlapping translates of a given centrally symmetric convex domain relative to its outer parallel domain of given outer radius is attained by a lattice packing in the Euclidean plane. This generalizes some earlier (classical) results. Sharp upper bounds are proved for the analogue problem on congruent circular disks in the spherical (resp., hyperbolic) plane and on congruent balls in Euclidean $3$-space.2024-01-01T03:15:39Z13 pages, 4 figuresPure and Applied Functional Analysis, Volume 10, Number 5 (November, 2025), Pages 1193-1205Károly BezdekZsolt Lángihttp://arxiv.org/abs/2401.00604v2SteinDreamer: Variance Reduction for Text-to-3D Score Distillation via Stein Identity2024-03-29T18:33:30ZScore distillation has emerged as one of the most prevalent approaches for text-to-3D asset synthesis. Essentially, score distillation updates 3D parameters by lifting and back-propagating scores averaged over different views. In this paper, we reveal that the gradient estimation in score distillation is inherent to high variance. Through the lens of variance reduction, the effectiveness of SDS and VSD can be interpreted as applications of various control variates to the Monte Carlo estimator of the distilled score. Motivated by this rethinking and based on Stein's identity, we propose a more general solution to reduce variance for score distillation, termed Stein Score Distillation (SSD). SSD incorporates control variates constructed by Stein identity, allowing for arbitrary baseline functions. This enables us to include flexible guidance priors and network architectures to explicitly optimize for variance reduction. In our experiments, the overall pipeline, dubbed SteinDreamer, is implemented by instantiating the control variate with a monocular depth estimator. The results suggest that SSD can effectively reduce the distillation variance and consistently improve visual quality for both object- and scene-level generation. Moreover, we demonstrate that SteinDreamer achieves faster convergence than existing methods due to more stable gradient updates.2023-12-31T23:04:25ZProject page: https://vita-group.github.io/SteinDreamer/Peihao WangZhiwen FanDejia XuDilin WangSreyas MohanForrest IandolaRakesh RanjanYilei LiQiang LiuZhangyang WangVikas Chandrahttp://arxiv.org/abs/2401.00607v1Ricci flows which terminate in cones2023-12-31T23:23:20ZWe prove that a complete solution to the Ricci flow on $M\times [-T, 0)$ which has quadratic curvature decay on some end of $M$ and converges locally smoothly to the end of a cone on that neighborhood as $t\nearrow 0$ must be a gradient shrinking soliton.2023-12-31T23:23:20Z52 pages, no figuresBrett Kotschwarhttp://arxiv.org/abs/2401.00619v2Collision energy dependence of source sizes for primary and secondary pions at NICA energies2024-06-06T04:03:03ZWe study the evolution of the two-pion correlation function parameters with collision energy in the context of relativistic heavy-ion collisions within the NICA energy range. To this end, we perform UrQMD simulations in the cascade mode to produce samples of pions from $5\times 10^6$ Bi+Bi collisions for each of the studied energies. The effects of the quantum-statistical correlations are introduced using the correlation afterburner code CRAB. We fit the correlation function using Gaussian, exponential and symmetric Lévy shapes and show that for all collision energies the latter provides the best fit. We separate the sample into pions coming from primary processes and pions originating from the decay of long-lived resonances, and show that the source size for the latter is significantly larger than for the former. The source size for the secondaries, is similar but in general larger than the size for the whole pion sample. To further characterize the pion source, we also simulate the effects of a non-ideal detector introducing a momentum smearing parameter, representing the minimum pair momentum and thus a maximum source size that can be resolved. The values of the correlation function intercept parameter are therefore modified from the values they attain for the perfect detector case. Using the core-halo picture of the source, we show that the values of the intercept parameter are influenced by the presence of a significant fraction of core pions coming from the decay of long-lived but slow-moving resonances. These findings serve as a benchmark to compare with future Monte Carlo studies that consider an Equation of State and thus allow for a phase transition within the studied energy domain.2024-01-01T00:18:24Z8 pages, 10 figures. Version accepted to appear in EPJ-AAlejandro AyalaSantiago Bernal-LangaricaIsabel DominguezIvonne MaldonadoMaria Elena Tejeda-Yeomanshttp://arxiv.org/abs/2401.00651v3IRWE: Inductive Random Walk for Joint Inference of Identity and Position Network Embedding2024-10-03T15:14:34ZNetwork embedding, which maps graphs to distributed representations, is a unified framework for various graph inference tasks. According to the topology properties (e.g., structural roles and community memberships of nodes) to be preserved, it can be categorized into the identity and position embedding. Most existing methods can only capture one type of property. Some approaches can support the inductive inference that generalizes the embedding model to new nodes or graphs but relies on the availability of attributes. Due to the complicated correlations between topology and attributes, it is unclear for some inductive methods which type of property they can capture. In this study, we explore a unified framework for the joint inductive inference of identity and position embeddings without attributes. An inductive random walk embedding (IRWE) method is proposed, which combines multiple attention units to handle the random walk (RW) on graph topology and simultaneously derives identity and position embeddings that are jointly optimized. We demonstrate that some RW statistics can characterize node identities and positions while supporting the inductive inference. Experiments validate the superior performance of IRWE over various baselines for the transductive and inductive inference of identity and position embeddings.2024-01-01T03:38:06ZAccepted by Transactions on Machine Learning Research (TMLR)Meng QinDit-Yan Yeunghttp://arxiv.org/abs/2401.00668v2The structure of the stellar halo of the Andromeda galaxy explored with the NB515 for Subaru/HSC. I.: New Insights on the stellar halo up to 120 kpc2025-01-02T08:20:58ZWe analyse the M31 halo and its substructure within a projected radius of 120 kpc using a combination of Subaru/HSC $\textit{NB515}$ and CFHT/MegaCam $\textit{g}$- \& $\textit{i}$-bands. We succeed in separating M31's halo stars from foreground contamination with $\sim$ 90 \% accuracy by using the surface gravity sensitive $\textit{NB515}$ filter. Based on the selected M31 halo stars, we discover three new substructures, which associate with the Giant Southern Stream (GSS) based on their photometric metallicity estimates. We also produce the distance and photometric metallicity estimates for the known substructures. While these quantities for the GSS are reproduced in our study, we find that the North-Western stream shows a steeper distance gradient than found in an earlier study, suggesting that it is likely to have formed in an orbit closer to the Milky Way. For two streams in the eastern halo (Stream C and D), we identify distance gradients that had not been resolved. Finally, we investigate the global halo photometric metallicity distribution and surface brightness profile using the $\textit{NB515}$-selected halo stars. We find that the surface brightness of the metal-poor and metal-rich halo populations, and the all population can be fitted to a power-law profile with an index of $α=-1.65\pm0.02$, $-2.82\pm0.01$, and $-2.44\pm0.01$, respectively. In contrast to the relative smoothness of the halo profile, its photometric metallicity distribution appears to be spatially non-uniform with nonmonotonic trends with radius, suggesting that the halo population had insufficient time to dynamically homogenize the accreted populations.2024-01-01T05:17:34Z24 pages, 26 figures, 5 tables, accepted for publication in MNRASItsuki OgamiMikito TanakaYutaka KomiyamaMasashi ChibaPuragra GuhathakurtaEvan N. KirbyRosemary F. G. WyseCarrie FilionKaroline M. GilbertIvanna EscalaMasao MoriTakanobu KiriharaMasayuki TanakaMiho N. IshigakiKohei HayashiMyung Gyoon LeeSanjib SharmaJason S. KaliraiRobert H. Luptonhttp://arxiv.org/abs/2401.00611v1A Compact Representation for Bayesian Neural Networks By Removing Permutation Symmetry2023-12-31T23:57:05ZBayesian neural networks (BNNs) are a principled approach to modeling predictive uncertainties in deep learning, which are important in safety-critical applications. Since exact Bayesian inference over the weights in a BNN is intractable, various approximate inference methods exist, among which sampling methods such as Hamiltonian Monte Carlo (HMC) are often considered the gold standard. While HMC provides high-quality samples, it lacks interpretable summary statistics because its sample mean and variance is meaningless in neural networks due to permutation symmetry. In this paper, we first show that the role of permutations can be meaningfully quantified by a number of transpositions metric. We then show that the recently proposed rebasin method allows us to summarize HMC samples into a compact representation that provides a meaningful explicit uncertainty estimate for each weight in a neural network, thus unifying sampling methods with variational inference. We show that this compact representation allows us to compare trained BNNs directly in weight space across sampling methods and variational inference, and to efficiently prune neural networks trained without explicit Bayesian frameworks by exploiting uncertainty estimates from HMC.2023-12-31T23:57:05ZAccepted at NeurIPS 2023 Workshop on Unifying Representations in Neural Models; 4 pages + appendixTim Z. XiaoWeiyang LiuRobert Bamlerhttp://arxiv.org/abs/2401.00616v3GD^2-NeRF: Generative Detail Compensation via GAN and Diffusion for One-shot Generalizable Neural Radiance Fields2024-03-29T11:27:32ZIn this paper, we focus on the One-shot Novel View Synthesis (O-NVS) task which targets synthesizing photo-realistic novel views given only one reference image per scene. Previous One-shot Generalizable Neural Radiance Fields (OG-NeRF) methods solve this task in an inference-time finetuning-free manner, yet suffer the blurry issue due to the encoder-only architecture that highly relies on the limited reference image. On the other hand, recent diffusion-based image-to-3d methods show vivid plausible results via distilling pre-trained 2D diffusion models into a 3D representation, yet require tedious per-scene optimization. Targeting these issues, we propose the GD$^2$-NeRF, a Generative Detail compensation framework via GAN and Diffusion that is both inference-time finetuning-free and with vivid plausible details. In detail, following a coarse-to-fine strategy, GD$^2$-NeRF is mainly composed of a One-stage Parallel Pipeline (OPP) and a 3D-consistent Detail Enhancer (Diff3DE). At the coarse stage, OPP first efficiently inserts the GAN model into the existing OG-NeRF pipeline for primarily relieving the blurry issue with in-distribution priors captured from the training dataset, achieving a good balance between sharpness (LPIPS, FID) and fidelity (PSNR, SSIM). Then, at the fine stage, Diff3DE further leverages the pre-trained image diffusion models to complement rich out-distribution details while maintaining decent 3D consistency. Extensive experiments on both the synthetic and real-world datasets show that GD$^2$-NeRF noticeably improves the details while without per-scene finetuning.2024-01-01T00:08:39ZSubmitted to JournalXiao PanZongxin YangShuai BaiYi Yanghttp://arxiv.org/abs/2401.00617v1Towards Improved Proxy-based Deep Metric Learning via Data-Augmented Domain Adaptation2024-01-01T00:10:58ZDeep Metric Learning (DML) plays an important role in modern computer vision research, where we learn a distance metric for a set of image representations. Recent DML techniques utilize the proxy to interact with the corresponding image samples in the embedding space. However, existing proxy-based DML methods focus on learning individual proxy-to-sample distance while the overall distribution of samples and proxies lacks attention. In this paper, we present a novel proxy-based DML framework that focuses on aligning the sample and proxy distributions to improve the efficiency of proxy-based DML losses. Specifically, we propose the Data-Augmented Domain Adaptation (DADA) method to adapt the domain gap between the group of samples and proxies. To the best of our knowledge, we are the first to leverage domain adaptation to boost the performance of proxy-based DML. We show that our method can be easily plugged into existing proxy-based DML losses. Our experiments on benchmarks, including the popular CUB-200-2011, CARS196, Stanford Online Products, and In-Shop Clothes Retrieval, show that our learning algorithm significantly improves the existing proxy losses and achieves superior results compared to the existing methods.2024-01-01T00:10:58ZAccepted by AAAI 2024Li RenChen ChenLiqiang WangKien Huahttp://arxiv.org/abs/2401.00634v2A scalable two-stage Bayesian approach accounting for exposure measurement error in environmental epidemiology2024-01-14T03:29:07ZAccounting for exposure measurement errors has been recognized as a crucial problem in environmental epidemiology for over two decades. Bayesian hierarchical models offer a coherent probabilistic framework for evaluating associations between environmental exposures and health effects, which take into account exposure measurement errors introduced by uncertainty in the estimated exposure as well as spatial misalignment between the exposure and health outcome data. While two-stage Bayesian analyses are often regarded as a good alternative to fully Bayesian analyses when joint estimation is not feasible, there has been minimal research on how to properly propagate uncertainty from the first-stage exposure model to the second-stage health model, especially in the case of a large number of participant locations along with spatially correlated exposures. We propose a scalable two-stage Bayesian approach, called a sparse multivariate normal (sparse MVN) prior approach, based on the Vecchia approximation for assessing associations between exposure and health outcomes in environmental epidemiology. We compare its performance with existing approaches through simulation. Our sparse MVN prior approach shows comparable performance with the fully Bayesian approach, which is a gold standard but is impossible to implement in some cases. We investigate the association between source-specific exposures and pollutant (nitrogen dioxide (NO$_2$))-specific exposures and birth outcomes for 2012 in Harris County, Texas, using several approaches, including the newly developed method.2024-01-01T02:07:26Z34 pages, 8 figuresChangwoo J. LeeElaine SymanskiAmal RammahDong Hun KangPhilip K. HopkeEun Sug Parkhttp://arxiv.org/abs/2401.00638v1A class of finite $p$-groups and the normalized unit groups of group algebras2024-01-01T02:30:52ZLet $p$ be a prime and $\mathbb{F}_p$ be a finite field of $p$ elements. Let $\mathbb{F}_pG$ denote the group algebra of the finite $p$-group $G$ over the field $\mathbb{F}_p$ and $V(\mathbb{F}_pG)$ denote the group of normalized units in $\mathbb{F}_pG$. Suppose that $G$ is a finite $p$-group given by a central extension of the form $$1\longrightarrow \mathbb{Z}_{p^n}\times \mathbb{Z}_{p^m} \longrightarrow G \longrightarrow \mathbb{Z}_p\times \cdots\times \mathbb{Z}_p \longrightarrow 1$$ and $G'\cong \mathbb{Z}_p$, $n, m\geq 1$ and $p$ is odd. In this paper, the structure of $G$ is determined. And the relations of $V(\mathbb{F}_pG)^{p^l}$ and $G^{p^l}$, $Ω_l(V(\mathbb{F}_pG))$ and $Ω_l(G)$ are given. Furthermore, there is a direct proof for $V(\mathbb{F}_pG)^p\bigcap G=G^p$.2024-01-01T02:30:52ZYulei WangHeguo Liuhttp://arxiv.org/abs/2401.00644v1DEWP: Deep Expansion Learning for Wind Power Forecasting2024-01-01T03:14:10ZWind is one kind of high-efficient, environmentally-friendly and cost-effective energy source. Wind power, as one of the largest renewable energy in the world, has been playing a more and more important role in supplying electricity. Though growing dramatically in recent years, the amount of generated wind power can be directly or latently affected by multiple uncertain factors, such as wind speed, wind direction, temperatures, etc. More importantly, there exist very complicated dependencies of the generated power on the latent composition of these multiple time-evolving variables, which are always ignored by existing works and thus largely hinder the prediction performances. To this end, we propose DEWP, a novel Deep Expansion learning for Wind Power forecasting framework to carefully model the complicated dependencies with adequate expressiveness. DEWP starts with a stack-by-stack architecture, where each stack is composed of (i) a variable expansion block that makes use of convolutional layers to capture dependencies among multiple variables; (ii) a time expansion block that applies Fourier series and backcast/forecast mechanism to learn temporal dependencies in sequential patterns. These two tailored blocks expand raw inputs into different latent feature spaces which can model different levels of dependencies of time-evolving sequential data. Moreover, we propose an inference block corresponding for each stack, which applies multi-head self-attentions to acquire attentive features and maps expanded latent representations into generated wind power. In addition, to make DEWP more expressive in handling deep neural architectures, we adapt doubly residue learning to process stack-by-stack outputs. Finally, we present extensive experiments in the real-world wind power forecasting application on two datasets from two different turbines to demonstrate the effectiveness of our approach.2024-01-01T03:14:10ZAccepted by TKDDWei FanYanjie FuShun ZhengJiang BianYuanchun ZhouHui Xionghttp://arxiv.org/abs/2401.00658v1Point Cloud in the Air2024-01-01T04:11:55ZAcquisition and processing of point clouds (PCs) is a crucial enabler for many emerging applications reliant on 3D spatial data, such as robot navigation, autonomous vehicles, and augmented reality. In most scenarios, PCs acquired by remote sensors must be transmitted to an edge server for fusion, segmentation, or inference. Wireless transmission of PCs not only puts on increased burden on the already congested wireless spectrum, but also confronts a unique set of challenges arising from the irregular and unstructured nature of PCs. In this paper, we meticulously delineate these challenges and offer a comprehensive examination of existing solutions while candidly acknowledging their inherent limitations. In response to these intricacies, we proffer four pragmatic solution frameworks, spanning advanced techniques, hybrid schemes, and distributed data aggregation approaches. In doing so, our goal is to chart a path toward efficient, reliable, and low-latency wireless PC transmission.2024-01-01T04:11:55ZYulin ShaoChenghong BianLi YangQianqian YangZhaoyang ZhangDeniz Gunduzhttp://arxiv.org/abs/2401.00662v1Enhancing Pre-trained ASR System Fine-tuning for Dysarthric Speech Recognition using Adversarial Data Augmentation2024-01-01T04:21:19ZAutomatic recognition of dysarthric speech remains a highly challenging task to date. Neuro-motor conditions and co-occurring physical disabilities create difficulty in large-scale data collection for ASR system development. Adapting SSL pre-trained ASR models to limited dysarthric speech via data-intensive parameter fine-tuning leads to poor generalization. To this end, this paper presents an extensive comparative study of various data augmentation approaches to improve the robustness of pre-trained ASR model fine-tuning to dysarthric speech. These include: a) conventional speaker-independent perturbation of impaired speech; b) speaker-dependent speed perturbation, or GAN-based adversarial perturbation of normal, control speech based on their time alignment against parallel dysarthric speech; c) novel Spectral basis GAN-based adversarial data augmentation operating on non-parallel data. Experiments conducted on the UASpeech corpus suggest GAN-based data augmentation consistently outperforms fine-tuned Wav2vec2.0 and HuBERT models using no data augmentation and speed perturbation across different data expansion operating points by statistically significant word error rate (WER) reductions up to 2.01% and 0.96% absolute (9.03% and 4.63% relative) respectively on the UASpeech test set of 16 dysarthric speakers. After cross-system outputs rescoring, the best system produced the lowest published WER of 16.53% (46.47% on very low intelligibility) on UASpeech.2024-01-01T04:21:19ZTo appear at IEEE ICASSP 2024Huimeng WangZengrui JinMengzhe GengShujie HuGuinan LiTianzi WangHaoning XuXunying Liuhttp://arxiv.org/abs/2401.00670v2Hybrid physics-informed metabolic cybergenetics: process rates augmented with machine-learning surrogates informed by flux balance analysis2024-03-25T14:37:40ZMetabolic cybergenetics is a promising concept that interfaces gene expression and cellular metabolism with computers for real-time dynamic metabolic control. The focus is on control at the transcriptional level, serving as a means to modulate intracellular metabolic fluxes. Recent strategies in this field have employed constraint-based dynamic models for process optimization, control, and estimation. However, this results in bilevel dynamic optimization problems, which pose considerable numerical and conceptual challenges. In this study, we present an alternative hybrid physics-informed dynamic modeling framework for metabolic cybergenetics, aimed at simplifying optimization, control, and estimation tasks. By utilizing machine-learning surrogates, our approach effectively embeds the physics of metabolic networks into the process rates of structurally simpler macro-kinetic models coupled with gene expression. These surrogates, informed by flux balance analysis, link the domains of manipulatable intracellular enzymes to metabolic exchange fluxes. This ensures that critical knowledge captured by the system's metabolic network is preserved. The resulting models can be integrated into metabolic cybergenetic schemes involving single-level optimizations. Additionally, the hybrid modeling approach maintains the number of system states at a necessary minimum, easing the burden of process monitoring and estimation. Our hybrid physics-informed metabolic cybergenetic framework is demonstrated using a computational case study on the optogenetically-assisted production of itaconate by $\textit{Escherichia coli}$.2024-01-01T05:20:43Z25 pages, 10 figures, journal submission (reviewed/accepted version)Sebastián Espinel-RíosJosé L. Avalos10.1021/acs.iecr.4c00001http://arxiv.org/abs/2401.00680v1Whittaker modules and hyperbolic Toda lattices2024-01-01T06:25:09ZLet $\sg$ be a complex finite-dimensional simple Lie algebra and let $\sg_l$ be the corresponding generalized Takiff algebra. This paper studies the affine variety $\ssf+\sb_l$ where $\ssf$ is similar to a principal nilpotent element of $\sg$ and $\sb_l$ is a subalgebra corresponding to the Borel subalgebra $\sb$ of $\sg$. Inspired by Kostant's work then we deal with two questions. One of them is to construct the Whittaker model for the $G_l$-invariants of symmetric algebra $S(\sg_l)$ where $G_l$ is the adjoint group of $\sg_l$ and $G_l$ acts on $S(\sg_l)$ by coadjoint action, and then to classify all nonsingular Whittaker modules over $\sg_l$. Another one is to describe the symplectic structure of the manifold $Z\subseteq\ssf+\sb_l$ of normalized Jacobi elements. Then the Hamiltonian corresponding to a fundamental invariant provides a class of hyperbolic Toda lattices. In particular, a simplest example describes the state of a dynamical system consisting of a positive mass particle and a negative mass particle.2024-01-01T06:25:09Z45 pagesLimeng Xiahttp://arxiv.org/abs/2401.00695v2Credible Teacher for Semi-Supervised Object Detection in Open Scene2024-01-03T02:33:49ZSemi-Supervised Object Detection (SSOD) has achieved resounding success by leveraging unlabeled data to improve detection performance. However, in Open Scene Semi-Supervised Object Detection (O-SSOD), unlabeled data may contains unknown objects not observed in the labeled data, which will increase uncertainty in the model's predictions for known objects. It is detrimental to the current methods that mainly rely on self-training, as more uncertainty leads to the lower localization and classification precision of pseudo labels. To this end, we propose Credible Teacher, an end-to-end framework. Credible Teacher adopts an interactive teaching mechanism using flexible labels to prevent uncertain pseudo labels from misleading the model and gradually reduces its uncertainty through the guidance of other credible pseudo labels. Empirical results have demonstrated our method effectively restrains the adverse effect caused by O-SSOD and significantly outperforms existing counterparts.2024-01-01T08:19:21ZAccpet by ICASSP 2024Jingyu ZhuangKuo WangLiang LinGuanbin Lihttp://arxiv.org/abs/2401.00698v1Large Language Models aren't all that you need2024-01-01T08:32:50ZThis paper describes the architecture and systems built towards solving the SemEval 2023 Task 2: MultiCoNER II (Multilingual Complex Named Entity Recognition) [1]. We evaluate two approaches (a) a traditional Conditional Random Fields model and (b) a Large Language Model (LLM) fine-tuned with a customized head and compare the two approaches. The novel ideas explored are: 1) Decaying auxiliary loss (with residual) - where we train the model on an auxiliary task of Coarse-Grained NER and include this task as a part of the loss function 2) Triplet token blending - where we explore ways of blending the embeddings of neighboring tokens in the final NER layer prior to prediction 3) Task-optimal heads - where we explore a variety of custom heads and learning rates for the final layer of the LLM. We also explore multiple LLMs including GPT-3 and experiment with a variety of dropout and other hyperparameter settings before arriving at our final model which achieves micro & macro f1 of 0.85/0.84 (on dev) and 0.67/0.61 on the test data . We show that while pre-trained LLMs, by themselves, bring about a large improvement in scores as compared to traditional models, we also demonstrate that tangible improvements to the Macro-F1 score can be made by augmenting the LLM with additional feature/loss/model engineering techniques described above.2024-01-01T08:32:50ZKiran Voderhobli HollaChaithanya KumarAryan Singhhttp://arxiv.org/abs/2401.00700v1An attempt to generate new bridge types from latent space of generative adversarial network2024-01-01T08:46:29ZTry to generate new bridge types using generative artificial intelligence technology. Symmetric structured image dataset of three-span beam bridge, arch bridge, cable-stayed bridge and suspension bridge are used . Based on Python programming language, TensorFlow and Keras deep learning platform framework , as well as Wasserstein loss function and Lipschitz constraints, generative adversarial network is constructed and trained. From the obtained low dimensional bridge-type latent space sampling, new bridge types with asymmetric structures can be generated. Generative adversarial network can create new bridge types by organically combining different structural components on the basis of human original bridge types. It has a certain degree of human original ability. Generative artificial intelligence technology can open up imagination space and inspire humanity.2024-01-01T08:46:29Z8 pages, 4 figuresHongjun Zhanghttp://arxiv.org/abs/2401.00679v2Second harmonic generation induced by gate voltage oscillation in few layer MnBi2Te42024-10-20T06:48:39ZNonlinear charge transport, such as nonreciprocal longitudinal resistance and nonlinear Hall effect, has attracted considerable interest in probing the symmetries and topological properties of new materials. Recent research has revealed significant nonreciprocal longitudinal resistance and nonlinear Hall effect in MnBi2Te4, an intrinsic magnetic topological insulator, induced by the quantum metric dipole. However, the inconsistent response with charge density and conflicting C3z symmetry requirement necessitate a thorough understanding of factors affecting the nonlinear transport measurement. This study uncovers an experimental factor leading to significant nonlinear transport signals in MnBi2Te4, attributed to gate voltage oscillation from the application of large alternating current. Additionally, a methodology is proposed to suppress this effect by individually grounding the voltage electrodes during second-harmonic measurements. The investigation underscores the critical importance of assessing the impact of gate voltage oscillation before determining the intrinsic nature of nonlinear transport in 2D material devices with an electrically connected operative gate electrode.2024-01-01T06:17:00Z16 pages, 4 figuresnpj Quantum Mater. 9, 79 (2024)Liangcai XuZichen LianYongchao WangXinlei HaoShuai YangYongqian WangChang LiuYang FengYayu WangJinsong Zhang10.1038/s41535-024-00694-8http://arxiv.org/abs/2401.00639v1Geometry Depth Consistency in RGBD Relative Pose Estimation2024-01-01T02:35:13ZRelative pose estimation for RGBD cameras is crucial in a number of applications. Previous approaches either rely on the RGB aspect of the images to estimate pose thus not fully making use of depth in the estimation process or estimate pose from the 3D cloud of points that each image produces, thus not making full use of RGB information. This paper shows that if one pair of correspondences is hypothesized from the RGB-based ranked-ordered correspondence list, then the space of remaining correspondences is restricted to corresponding pairs of curves nested around the hypothesized correspondence, implicitly capturing depth consistency. This simple Geometric Depth Constraint (GDC) significantly reduces potential matches. In effect this becomes a filter on possible correspondences that helps reduce the number of outliers and thus expedites RANSAC significantly. As such, the same budget of time allows for more RANSAC iterations and therefore additional robustness and a significant speedup. In addition, the paper proposed a Nested RANSAC approach that also speeds up the process, as shown through experiments on TUM, ICL-NUIM, and RGBD Scenes v2 datasets.2024-01-01T02:35:13ZSourav KumarChiang-Heng ChienBenjamin Kimiahttp://arxiv.org/abs/2401.00659v4Distinctiveness Maximization in Datasets Assemblage2025-02-27T11:38:18ZIn this paper, given a user's query set and budget, we aim to use the limited budget to help users assemble a set of datasets that can enrich a base dataset by introducing the maximum number of distinct tuples (i.e., maximizing distinctiveness). We prove this problem to be NP-hard. A greedy algorithm using exact distinctiveness computation attains an approximation ratio of (1-1/e)/2, but it lacks efficiency and scalability due to its frequent computation of the exact distinctiveness marginal gain of any candidate dataset for selection. This requires scanning through every tuple in candidate datasets and thus is unaffordable in practice. To overcome this limitation, we propose an efficient machine learning (ML)-based method for estimating the distinctiveness marginal gain of any candidate dataset. This effectively eliminates the need to test each tuple individually. Estimating the distinctiveness marginal gain of a dataset involves estimating the number of distinct tuples in the tuple sets returned by each query in a query set across multiple datasets. This can be viewed as the cardinality estimation for a query set on a set of datasets, and the proposed method is the first to tackle this cardinality estimation problem. This is a significant advancement over prior methods that were limited to single-query cardinality estimation on a single dataset and struggled with identifying overlaps among tuple sets returned by each query in a query set across multiple datasets. Extensive experiments using five real-world data pools demonstrate that our algorithm, which utilizes ML-based distinctiveness estimation, outperforms all relevant baselines in effectiveness, efficiency, and scalability. A case study on two downstream ML tasks also highlights its potential to find datasets with more useful tuples to enhance the performance of ML tasks.2024-01-01T04:14:00ZThis is a technical report of an accepted WWW'25 workTingting WangShixun HuangZhifeng BaoJ. Shane CulpepperVolkan DedeogluReza Arabloueihttp://arxiv.org/abs/2401.00683v3Asymptotically Optimal Sequence Sets With Low/Zero Ambiguity Zone Properties2025-03-19T01:20:43ZSequences with low/zero ambiguity zone (LAZ/ZAZ) properties are useful in modern communication and radar systems operating over mobile environments. This paper first presents a new family of ZAZ sequence sets motivated by the ``modulating'' zero correlation zone (ZCZ) sequences which were first proposed by Popovic and Mauritz. We then introduce a second family of ZAZ sequence sets with comb-like spectrum, whereby the local Doppler resilience is guaranteed by their inherent spectral nulls in the frequency domain. Finally, LAZ sequence sets are obtained by exploiting their connection with a novel class of mapping functions. These proposed unimodular ZAZ and LAZ sequence sets are cyclically distinct and asymptotically optimal with respect to the existing theoretical bounds on ambiguity functions.2024-01-01T06:58:31ZLiying TianXiaoshi SongZilong LiuYubo Lihttp://arxiv.org/abs/2401.00635v3Optimization of deterministic photonic graph state generation via local operations2024-12-06T19:41:08ZRealizing photonic graph states, crucial in various quantum protocols, is challenging due to the absence of deterministic entangling gates in linear optics. To address this, emitter qubits have been leveraged to establish and transfer the entanglement to photons. We introduce an optimization method for such protocols based on the local Clifford equivalency of states and the graph theoretical correlations of the generation cost parameters. Employing this method, we achieve a 50% reduction in use of the 2-qubit gates for generation of the arbitrary large repeater graph states and similar significant reductions in the total gate count for generation of random dense graphs.2024-01-01T02:11:49Z12 pages, 15 figuresPhys. Rev. A 110, (2024) 052605Sobhan GhanbariJie LinBenjamin MacLellanLuc RobichaudPiotr RoztockiHoi-Kwong Lo10.1103/PhysRevA.110.052605http://arxiv.org/abs/2401.00614v3On Nontrivial Winning and Losing Parameters of Schmidt Games2025-02-20T09:11:14ZIn this paper we study the classical Schmidt game on two families of sets: one related to frequencies of digits in base-$2$ expansions, and one connected to the set of the badly approximable numbers. Namely, we describe some nontrivial winning and losing parameters $(α, β)$ for these sets.2024-01-01T00:01:21ZSome issues are corrected // 21 pages, 7 figuresResults Math 80, 236 (2025)Vasiliy NeckrasovEric Zhan10.1007/s00025-025-02547-7http://arxiv.org/abs/2401.00626v2Complex continued fractions, Kleinian and extremal theory for cusp excursions2025-10-13T10:19:23ZFor the each of the five Euclidean rings of complex quadratic integers, we consider a complex continued fraction algorithm with digits in the ring. We show for each algorithm that the maximal digit obeys a Fréchet distribution. We use this to find a limiting distribution for cusp excursions on Bianchi orbifolds associated with the aforementioned rings of quadratic integers.2024-01-01T01:14:15Z19 pages, 4 figures Fixed typos. Converted appendix to new section and improved presentation of proofs. Added Theorem 1.3 (an application to Diophantine Approximation) along with a proof in Subsection 6.3. Changed title to more accurately reflect contents of preprintAlexander BaumgartnerMark Pollicotthttp://arxiv.org/abs/2401.00606v1Backward propagation of warped product structures and asymptotically conical shrinkers2023-12-31T23:17:34ZWe establish sufficient conditions which ensure that a locally-warped product structure propagates backward in time under the Ricci flow. As an application, we prove that if an asymptotically conical gradient shrinking soliton is asymptotic to a cone whose cross-section is a product of Einstein manifolds, the soliton must itself be a multiply-warped product over the same manifolds.2023-12-31T23:17:34Z56 pages, no figuresBrett Kotschwarhttp://arxiv.org/abs/2401.00621v2Multiplicity of normalized solutions for the fractional Schrödinger equation with potentials2024-01-22T10:59:48ZWe get multiplicity of normalized solutions for the fractional Schrödinger equation
$$
(-Δ)^su+V(\varepsilon x)u=λu+h(\varepsilon x)f(u)\quad \mbox{in $\mathbb{R}^N$},
\qquad\int_{\mathbb{R}^N}|u|^2dx=a,
$$ where $(-Δ)^s$ is the fractional Laplacian, $s\in(0,1)$, $a,\varepsilon>0$, $λ\in\mathbb{R}$ is an unknown parameter that appears as a Lagrange multiplier, $V,h:\mathbb{R}^N\rightarrow[0,+\infty)$ are bounded and continuous, and $f$ is continuous function with $L^2$-subcritical growth. We prove that the numbers of normalized solutions are at least the numbers of global maximum points of $h$ when $\varepsilon$ is small enough.2024-01-01T00:39:08Z19 pages, revised versionXue ZhangMarco SquassinaJianjun Zhanghttp://arxiv.org/abs/2401.00627v1Student and AI responses to physics problems examined through the lenses of sensemaking and mechanistic reasoning2024-01-01T01:16:39ZSeveral reports in education have called for transforming physics learning environments by promoting sensemaking of real-world scenarios in light of curricular ideas. Recent advancements in Generative-Artificial Intelligence has garnered increasing traction in educators' community by virtue of its potential in transforming STEM learning. In this exploratory study, we adopt a mixed-methods approach in comparatively examining student- and AI-generated responses to two different formats of a physics problem through the cognitive lenses of sensemaking and mechanistic reasoning. The student data is derived from think-aloud interviews of introductory students and the AI data comes from ChatGPT's solutions collected using Zero shot approach. The results highlight AI responses to evidence most features of the two processes through well-structured solutions and student responses to effectively leverage representations in their solutions through iterative refinement of arguments. In other words, while AI responses reflect how physics is talked about, the student responses reflect how physics is practiced. Implications of these results in light of development and deployment of AI systems in physics pedagogy are discussed.2024-01-01T01:16:39ZAmogh SirnoorkarDean ZollmanJames T. LavertyAlejandra J. MaganaSanjay RebelloLynn A. Bryanhttp://arxiv.org/abs/2401.00633v1On Discprecncies between Perturbation Evaluations of Graph Neural Network Attributions2024-01-01T02:03:35ZNeural networks are increasingly finding their way into the realm of graphs and modeling relationships between features. Concurrently graph neural network explanation approaches are being invented to uncover relationships between the nodes of the graphs. However, there is a disparity between the existing attribution methods, and it is unclear which attribution to trust. Therefore research has introduced evaluation experiments that assess them from different perspectives. In this work, we assess attribution methods from a perspective not previously explored in the graph domain: retraining. The core idea is to retrain the network on important (or not important) relationships as identified by the attributions and evaluate how networks can generalize based on these relationships. We reformulate the retraining framework to sidestep issues lurking in the previous formulation and propose guidelines for correct analysis. We run our analysis on four state-of-the-art GNN attribution methods and five synthetic and real-world graph classification datasets. The analysis reveals that attributions perform variably depending on the dataset and the network. Most importantly, we observe that the famous GNNExplainer performs similarly to an arbitrary designation of edge importance. The study concludes that the retraining evaluation cannot be used as a generalized benchmark and recommends it as a toolset to evaluate attributions on a specifically addressed network, dataset, and sparsity.2024-01-01T02:03:35ZRazieh RezaeiAlireza DizajiAshkan KhakzarAnees KaziNassir NavabDaniel Rueckerthttp://arxiv.org/abs/2401.00648v1An example for Kuznetsov-Shinder conjecture2024-01-01T03:31:46ZWe give an example for the Kuznetsov-Shinder conjecture, with infinitely many non-isomorphic D-equivalent and L-equivalent varieties.2024-01-01T03:31:46Z5 pages. This work was prepared for talk at 38th Annual Conference of Ramanujan Mathematical Society, 2023Tanya Kaushal Srivastavahttp://arxiv.org/abs/2401.00652v1From Covert Hiding to Visual Editing: Robust Generative Video Steganography2024-01-01T03:40:07ZTraditional video steganography methods are based on modifying the covert space for embedding, whereas we propose an innovative approach that embeds secret message within semantic feature for steganography during the video editing process. Although existing traditional video steganography methods display a certain level of security and embedding capacity, they lack adequate robustness against common distortions in online social networks (OSNs). In this paper, we introduce an end-to-end robust generative video steganography network (RoGVS), which achieves visual editing by modifying semantic feature of videos to embed secret message. We employ face-swapping scenario to showcase the visual editing effects. We first design a secret message embedding module to adaptively hide secret message into the semantic feature of videos. Extensive experiments display that the proposed RoGVS method applied to facial video datasets demonstrate its superiority over existing video and image steganography techniques in terms of both robustness and capacity.2024-01-01T03:40:07ZUnder ReviewXueying MaoXiaoxiao HuWanli PengZhenliang GanQichao YingZhenxing QianSheng LiXinpeng Zhanghttp://arxiv.org/abs/2401.00653v1PROMPT-IML: Image Manipulation Localization with Pre-trained Foundation Models Through Prompt Tuning2024-01-01T03:45:07ZDeceptive images can be shared in seconds with social networking services, posing substantial risks. Tampering traces, such as boundary artifacts and high-frequency information, have been significantly emphasized by massive networks in the Image Manipulation Localization (IML) field. However, they are prone to image post-processing operations, which limit the generalization and robustness of existing methods. We present a novel Prompt-IML framework. We observe that humans tend to discern the authenticity of an image based on both semantic and high-frequency information, inspired by which, the proposed framework leverages rich semantic knowledge from pre-trained visual foundation models to assist IML. We are the first to design a framework that utilizes visual foundation models specially for the IML task. Moreover, we design a Feature Alignment and Fusion module to align and fuse features of semantic features with high-frequency features, which aims at locating tampered regions from multiple perspectives. Experimental results demonstrate that our model can achieve better performance on eight typical fake image datasets and outstanding robustness.2024-01-01T03:45:07ZUnder ReviewXuntao LiuYuzhou YangQichao YingZhenxing QianXinpeng ZhangSheng Lihttp://arxiv.org/abs/2401.00655v1The minimal periodic solutions for superquadratic autonomous Hamiltonian systems without the Palais-Smale condition2024-01-01T03:53:32ZIn this paper, we prove the existence of periodic solutions with any prescribed minimal period $T>0$ for even second order Hamiltonian systems and convex first order Hamiltonian systems under the weak Nehari condition instead of Ambrosetti-Rabinowitz's. To this end, we shall develop the method of Nehari manifold to directly deal with a frequently occurring problem where the Nehari manifold is not a manifold.2024-01-01T03:53:32ZYuming XiaoGaosheng Zhuhttp://arxiv.org/abs/2401.00656v1Scalable iterative data-adaptive RKHS regularization2024-01-01T03:58:17ZWe present iDARR, a scalable iterative Data-Adaptive RKHS Regularization method, for solving ill-posed linear inverse problems. The method searches for solutions in subspaces where the true solution can be identified, with the data-adaptive RKHS penalizing the spaces of small singular values. At the core of the method is a new generalized Golub-Kahan bidiagonalization procedure that recursively constructs orthonormal bases for a sequence of RKHS-restricted Krylov subspaces. The method is scalable with a complexity of $O(kmn)$ for $m$-by-$n$ matrices with $k$ denoting the iteration numbers. Numerical tests on the Fredholm integral equation and 2D image deblurring show that it outperforms the widely used $L^2$ and $l^2$ norms, producing stable accurate solutions consistently converging when the noise level decays.2024-01-01T03:58:17ZHaibo LiJinchao FengFei Luhttp://arxiv.org/abs/2401.00657v1Optimizing ADMM and Over-Relaxed ADMM Parameters for Linear Quadratic Problems2024-01-01T04:01:40ZThe Alternating Direction Method of Multipliers (ADMM) has gained significant attention across a broad spectrum of machine learning applications. Incorporating the over-relaxation technique shows potential for enhancing the convergence rate of ADMM. However, determining optimal algorithmic parameters, including both the associated penalty and relaxation parameters, often relies on empirical approaches tailored to specific problem domains and contextual scenarios. Incorrect parameter selection can significantly hinder ADMM's convergence rate. To address this challenge, in this paper we first propose a general approach to optimize the value of penalty parameter, followed by a novel closed-form formula to compute the optimal relaxation parameter in the context of linear quadratic problems (LQPs). We then experimentally validate our parameter selection methods through random instantiations and diverse imaging applications, encompassing diffeomorphic image registration, image deblurring, and MRI reconstruction.2024-01-01T04:01:40ZAccepted to AAAI 2024Jintao SongWenqi LuYunwen LeiYuchao TangZhenkuan PanJinming Duanhttp://arxiv.org/abs/2401.00660v2Heavy quark structure functions from unifying the color dipole picture and double asymptotic scaling approaches2024-02-10T05:20:11ZWe present an analysis of the heavy quark structure functions from the $k_{t}$ factorization scheme, using unifying the color dipole picture and double asymptotic scaling approaches at small $x$. The gluon distribution is obtained from the Golec-Biernat-W$\ddot{\mathrm{u}}$sthoff (GBW) and Bartels, Golec-Biernat and Kowalski (BGK )models. The main elements are based on the color dipole picture (CDP) and the generalized double asymptotic scaling (DAS) approach for usual parton distribution functions (PDFs). The comparisons with the HERA data are made and predictions for the proposed LHeC and FCC-he colliders are also provided in a wide range of the transverse separation $r$. In particular, the ratio $R^{h}=F_{L}^{h}/F_{2}^{h}, h=c,b,t$ is well described by the dipole models and is sensitive to the collider energies from HERA until FCC-he. We derive correlated bounds on the ratio $F^{c}_{2}/F_{2}$ and $F^{b}_{2}/F_{2}$ and compared them with the BGK and IP-sat models. The uncertainties are due to the renormalization and factorization scales at large and low $r$ values. The Sudakov form factor into the heavy quark structure functions is incorporated and the results are considered, which are dependent on the hard scale in a wide range of the transverse separation $r$.2024-01-01T04:20:08ZPhysical Review D 109, 054012 (2024)G. R. Boroun10.1103/PhysRevD.109.054012http://arxiv.org/abs/2401.00663v11st Place Solution for 5th LSVOS Challenge: Referring Video Object Segmentation2024-01-01T04:24:48ZThe recent transformer-based models have dominated the Referring Video Object Segmentation (RVOS) task due to the superior performance. Most prior works adopt unified DETR framework to generate segmentation masks in query-to-instance manner. In this work, we integrate strengths of that leading RVOS models to build up an effective paradigm. We first obtain binary mask sequences from the RVOS models. To improve the consistency and quality of masks, we propose Two-Stage Multi-Model Fusion strategy. Each stage rationally ensembles RVOS models based on framework design as well as training strategy, and leverages different video object segmentation (VOS) models to enhance mask coherence by object propagation mechanism. Our method achieves 75.7% J&F on Ref-Youtube-VOS validation set and 70% J&F on test set, which ranks 1st place on 5th Large-scale Video Object Segmentation Challenge (ICCV 2023) track 3. Code is available at https://github.com/RobertLuo1/iccv2023_RVOS_Challenge.2024-01-01T04:24:48ZZhuoyan LuoYicheng XiaoYong LiuYitong WangYansong TangXiu LiYujiu Yanghttp://arxiv.org/abs/2401.00684v1A Temporal Filter to Extract Doped Conducting Polymer Information Features from an Electronic Nose2024-01-01T07:04:20ZIdentifying relevant machine-learning features for multi-sensing platforms is both an applicative limitation to recognize environments and a necessity to interpret the physical relevance of transducers' complementarity in their information processing. Particularly for long acquisitions, feature extraction must be fully automatized without human intervention and resilient to perturbations without increasing significantly the computational cost of a classifier. In this study, we investigate on the relative resistance and current modulation of a 24-dimensional conductimetric electronic nose, which uses the exponential moving average as a floating reference in a low-cost information descriptor for environment recognition. In particular, we identified that depending on the structure of a linear classifier, the 'modema' descriptor is optimized for different material sensing elements' contributions to classify information patterns. The low-pass filtering optimization leads to opposite behaviors between unsupervised and supervised learning: the latter one favors longer integration of the reference, allowing to recognize five different classes over 90%, while the first one prefers using the latest events as its reference to clusterize patterns by environment nature. Its electronic implementation shall greatly diminish the computational requirements of conductimetric electronic noses for on-board environment recognition without human supervision.2024-01-01T07:04:20ZWiem Haj AmmarAicha BoujnahAntoine BaronAimen BoubakerAdel KalboussiKamal LmimouniSebastien Pecqueurhttp://arxiv.org/abs/2401.00685v2Communication-Efficient Federated Learning for LEO Satellite Networks Integrated with HAPs Using Hybrid NOMA-OFDM2024-02-16T09:21:29ZSpace AI has become increasingly important and sometimes even necessary for government, businesses, and society. An active research topic under this mission is integrating federated learning (FL) with satellite communications (SatCom) so that numerous low Earth orbit (LEO) satellites can collaboratively train a machine learning model. However, the special communication environment of SatCom leads to a very slow FL training process up to days and weeks. This paper proposes NomaFedHAP, a novel FL-SatCom approach tailored to LEO satellites, that (1) utilizes high-altitude platforms (HAPs) as distributed parameter servers (PS) to enhance satellite visibility, and (2) introduces non-orthogonal multiple access (NOMA) into LEO to enable fast and bandwidth-efficient model transmissions. In addition, NomaFedHAP includes (3) a new communication topology that exploits HAPs to bridge satellites among different orbits to mitigate the Doppler shift, and (4) a new FL model aggregation scheme that optimally balances models between different orbits and shells. Moreover, we (5) derive a closed-form expression of the outage probability for satellites in near and far shells, as well as for the entire system. Our extensive simulations have validated the mathematical analysis and demonstrated the superior performance of NomaFedHAP in achieving fast and efficient FL model convergence with high accuracy as compared to the state-of-the-art.2024-01-01T07:07:27ZMohamed ElmahallawyTie LuoKhaled Ramadanhttp://arxiv.org/abs/2401.00625v4Beyond Efficiency: A Systematic Survey of Resource-Efficient Large Language Models2024-12-29T17:38:32ZThe burgeoning field of Large Language Models (LLMs), exemplified by sophisticated models like OpenAI's ChatGPT, represents a significant advancement in artificial intelligence. These models, however, bring forth substantial challenges in the high consumption of computational, memory, energy, and financial resources, especially in environments with limited resource capabilities. This survey aims to systematically address these challenges by reviewing a broad spectrum of techniques designed to enhance the resource efficiency of LLMs. We categorize methods based on their optimization focus: computational, memory, energy, financial, and network resources and their applicability across various stages of an LLM's lifecycle, including architecture design, pretraining, finetuning, and system design. Additionally, the survey introduces a nuanced categorization of resource efficiency techniques by their specific resource types, which uncovers the intricate relationships and mappings between various resources and corresponding optimization techniques. A standardized set of evaluation metrics and datasets is also presented to facilitate consistent and fair comparisons across different models and techniques. By offering a comprehensive overview of the current sota and identifying open research avenues, this survey serves as a foundational reference for researchers and practitioners, aiding them in developing more sustainable and efficient LLMs in a rapidly evolving landscape.2024-01-01T01:12:42ZGitHub repo: https://github.com/tiingweii-shii/Awesome-Resource-Efficient-LLM-PapersGuangji BaiZheng ChaiChen LingShiyu WangJiaying LuNan ZhangTingwei ShiZiyang YuMengdan ZhuYifei ZhangXinyuan SongCarl YangYue ChengLiang Zhaohttp://arxiv.org/abs/2401.00665v4An algorithm for estimating the crossing number of dense graphs, and continuous analogs of the crossing and rectilinear crossing numbers2025-01-10T09:36:56ZWe present a deterministic $n^{2+o(1)}$-time algorithm that approximates the crossing number of any graph $G$ of order $n$ up to an additive error of $o(n^4)$. We also provide a randomized polynomial-time algorithm that constructs a drawing of $G$ with $\text{cr}(G)+o(n^4)$ crossings. These results yield a $1+o(1)$ approximation algorithm for the crossing number of dense graphs. Our work complements a paper of Fox, Pach and Súk, who obtained similar results for the rectilinear crossing number.
The results of Fox, Pach and Súk and in this paper imply that the (normalized) crossing and rectilinear crossing numbers are estimable parameters. Motivated by this, we introduce two graphon parameters, the \textit{crossing density} and the \textit{rectilinear crossing density}, and we prove that, in a precise sense, these are the correct continuous analogs of the crossing and rectilinear crossing numbers of graphs.2024-01-01T04:47:32Z24 pages, 4 figuresOriol Solé-Pihttp://arxiv.org/abs/2401.00673v1Large deviation principle for slow-fast rough differential equations via controlled rough paths2024-01-01T05:34:31ZWe prove a large deviation principle for the slow-fast rough differential equations under the controlled rough path framework. The driver rough paths are lifted from the mixed fractional Brownian motion with Hurst parameter $H\in (1/3,1/2)$. Our approach is based on the continuity of the solution mapping and the variational framework for mixed fractional Brownian motion. By utilizing the variational representation, our problem is transformed into a qualitative property of the controlled system. In particular, the fast rough differential equation coincides with Itô SDE almost surely, which possesses a unique invariant probability measure with frozen slow component. We then demonstrate the weak convergence of the controlled slow component by averaging with respect to the invariant measure of the fast equation and exploiting the continuity of the solution mapping.2024-01-01T05:34:31ZXiaoyu YangYong Xu10.1017/prm.2025.4http://arxiv.org/abs/2401.00613v1Extracting spectra in the shell model Monte Carlo method using imaginary-time correlation matrices2023-12-31T23:59:21ZConventional diagonalization methods to calculate nuclear energy levels in the framework of the configuration-interaction (CI) shell model approach are prohibited in very large model spaces. The shell model Monte Carlo (SMMC) is a powerful technique for calculating thermal and ground-state observables of nuclei in very large model spaces, but it is challenging to extract nuclear spectra in this approach. We present a novel method to extract low-lying energy levels for given values of a set of good quantum numbers such as spin and parity. The method is based on imaginary-time one-body density correlation matrices that satisfy asymptotically a generalized eigenvalue problem. We validate the method in a light nucleus that allows comparison with exact diagonalization results of the CI shell model Hamiltonian. The method is applicable to other finite-size quantum many-body systems that can be described within a CI shell model approach.2023-12-31T23:59:21Z5 pages, 2 figuresPhys. Rev. Lett. 133, 182501 (2024)Y. AlhassidM. Bonett-MatizC. N. GilbrethS. Vartakhttp://arxiv.org/abs/2401.00608v5Reviving the Context: Camera Trap Species Classification as Link Prediction on Multimodal Knowledge Graphs2024-08-24T16:13:24ZCamera traps are important tools in animal ecology for biodiversity monitoring and conservation. However, their practical application is limited by issues such as poor generalization to new and unseen locations. Images are typically associated with diverse forms of context, which may exist in different modalities. In this work, we exploit the structured context linked to camera trap images to boost out-of-distribution generalization for species classification tasks in camera traps. For instance, a picture of a wild animal could be linked to details about the time and place it was captured, as well as structured biological knowledge about the animal species. While often overlooked by existing studies, incorporating such context offers several potential benefits for better image understanding, such as addressing data scarcity and enhancing generalization. However, effectively incorporating such heterogeneous context into the visual domain is a challenging problem. To address this, we propose a novel framework that transforms species classification as link prediction in a multimodal knowledge graph (KG). This framework enables the seamless integration of diverse multimodal contexts for visual recognition. We apply this framework for out-of-distribution species classification on the iWildCam2020-WILDS and Snapshot Mountain Zebra datasets and achieve competitive performance with state-of-the-art approaches. Furthermore, our framework enhances sample efficiency for recognizing under-represented species.2023-12-31T23:32:03Z12 pages, 5 figuresVardaan PahujaWeidi LuoYu GuCheng-Hao TuHong-You ChenTanya Berger-WolfCharles StewartSong GaoWei-Lun ChaoYu Su10.1145/3627673.3679545http://arxiv.org/abs/2401.00649v2Linear Model and Extensions2025-06-19T05:34:30ZI developed the lecture notes based on my ``Linear Model'' course at the University of California, Berkeley over the past ten years. This book provides an intermediate-level introduction to the linear model. It balances rigorous proofs and heuristic arguments. This book provides R code to replicate all simulation studies and case studies.2024-01-01T03:34:17ZMore polished version with R code on Harvard DataversePeng Dinghttp://arxiv.org/abs/2401.00664v7Metric Entropy-Free Sample Complexity Bounds for Sample Average Approximation in Convex Stochastic Programming2026-03-02T18:57:52ZThis paper studies sample average approximation (SAA) in solving convex or strongly convex stochastic programming (SP) problems. In estimating SAA's sample efficiency, the state-of-the-art sample complexity bounds entail metric entropy terms (such as the logarithm of the feasible region's covering number), which often grow polynomially with problem dimensionality. While it has been shown that metric entropy-free complexity rates are attainable under a uniform Lipschitz condition, such an assumption can be overly critical for many important SP problem settings. In response, this paper presents metric entropy-free sample complexity bounds for the SAA under standard SP assumptions} -- in the absence of the uniform Lipschitz condition. For a $d$-dimensional problem, the new results often lead to an $O(d)$-improvement in the complexity rate compared with the state-of-the-art. From the newly established complexity bounds, an important revelation is that SAA and the canonical stochastic mirror descent (SMD) method, two mainstream solution approaches to SP, entail almost identical rates of sample efficiency, lifting a theoretical discrepancy of SAA from SMD also by a factor of $O(d)$. Furthermore, this paper explores non-Lipschitzian scenarios where SAA maintains provable efficacy but the corresponding results for SMD remain mostly unexplored, indicating the potential of SAA's better applicability in some irregular settings. The results of our numerical experiments align with our theoretical findings.2024-01-01T04:35:53ZHongcheng LiuJindong Tonghttp://arxiv.org/abs/2401.00618v4Changes-in-Changes for Ordered Choice Models with Underreporting2026-04-18T12:33:02ZWe develop a Difference-in-Differences framework for discrete, ordered outcomes subject to underreporting. Such outcomes commonly arise in self-reported surveys on socially undesirable or stigmatized behaviors, where respondents may conceal their true behavior. For a discrete Changes-in-Changes model that is shown to admit an equivalent threshold-crossing representation, we derive nonparametric bounds for the counterfactual and factual outcome distributions as well as for the associated quantile treatment effects when outcomes are underreported. These bounds are shown to be sharp uniformly across outcome levels under additional support conditions, and we propose suitable estimation and bootstrap inference procedures. In an extension, we also consider a semiparametric underreporting model that allows to point identify and estimate distributional treatment effects. As an application, we investigate the impact of recreational marijuana legalization on the consumption behavior of 8th-grade students in several U.S. states.2024-01-01T00:12:56ZDaniel GutknechtCenchen Liuhttp://arxiv.org/abs/2401.00609v1A Survey of Personality, Persona, and Profile in Conversational Agents and Chatbots2023-12-31T23:41:41ZWe present a review of personality in neural conversational agents (CAs), also called chatbots. First, we define Personality, Persona, and Profile. We explain all personality schemes which have been used in CAs, and list models under the scheme(s) which they use. Second we describe 21 datasets which have been developed in recent CA personality research. Third, we define the methods used to embody personality in a CA, and review recent models using them. Fourth, we survey some relevant reviews on CAs, personality, and related topics. Finally, we draw conclusions and identify some research challenges for this important emerging field.2023-12-31T23:41:41Z25 pages, 6 tables, 207 referencesRichard Sutcliffehttp://arxiv.org/abs/2401.00628v1On the 2D Yang-Mills/Hurwitz Correspondence2024-01-01T01:26:53ZIn this paper, we show that in the large $N$ limit two-dimensional Yang-Mills theory with $U(N)$ gauge group becomes mixed Hurwitz theory, in the sense that the $1/N$ expansion of the chiral partition function receives contributions from both classical and monotone Hurwitz theory for all but finitely many compact orientable spacetimes.2024-01-01T01:26:53Z35 pages, 2 figuresJonathan Novakhttp://arxiv.org/abs/2401.00630v3An almost linear time algorithm testing whether the Markoff graph modulo $p$ is connected2024-01-26T04:26:34ZThe Markoff graph modulo $p$ is known to be connected for all but finitely many primes $p$ (see Eddy, Fuchs, Litman, Martin, Tripeny, and Vanyo [arxiv:2308.07579]), and it is conjectured that these graphs are connected for all primes. In this paper, we provide an algorithmic realization of the process introduced by Bourgain, Gamburd, and Sarnak [arxiv:1607.01530] to test whether the Markoff graph modulo $p$ is connected for arbitrary primes. Our algorithm runs in $o(p^{1 + ε})$ time for every $ε> 0$. We demonstrate this algorithm by confirming that the Markoff graph modulo $p$ is connected for all primes less than one million.2024-01-01T01:54:01ZVersion 3: Present more data and fix typosColby Austin Brownhttp://arxiv.org/abs/2401.00632v1TBDD: A New Trust-based, DRL-driven Framework for Blockchain Sharding in IoT2024-01-01T01:57:28ZIntegrating sharded blockchain with IoT presents a solution for trust issues and optimized data flow. Sharding boosts blockchain scalability by dividing its nodes into parallel shards, yet it's vulnerable to the $1\%$ attacks where dishonest nodes target a shard to corrupt the entire blockchain. Balancing security with scalability is pivotal for such systems. Deep Reinforcement Learning (DRL) adeptly handles dynamic, complex systems and multi-dimensional optimization. This paper introduces a Trust-based and DRL-driven (\textsc{TbDd}) framework, crafted to counter shard collusion risks and dynamically adjust node allocation, enhancing throughput while maintaining network security. With a comprehensive trust evaluation mechanism, \textsc{TbDd} discerns node types and performs targeted resharding against potential threats. The model maximizes tolerance for dishonest nodes, optimizes node movement frequency, ensures even node distribution in shards, and balances sharding risks. Rigorous evaluations prove \textsc{TbDd}'s superiority over conventional random-, community-, and trust-based sharding methods in shard risk equilibrium and reducing cross-shard transactions.2024-01-01T01:57:28ZZixu ZhangGuangsheng YuCaijun SunXu WangYing WangMing ZhangWei NiRen Ping LiuAndrew ReevesNektarios Georgalashttp://arxiv.org/abs/2401.00643v1Spectral action and heat kernel trace for Ricci flat manifolds from stochastic flow over second quantized $L^2$-differential forms2024-01-01T03:08:36ZA quantum stochastic differential equation (qsde) on Fock space over $L^2$ differential 1-forms is given from the small "time" flow of which the trace of the connection Laplacian heat kernel for the spinor endomorphism bundle can be computed over any compact Ricci-flat Riemannian manifold. The existence of the stochastic flow is established by adapting the construction from [14]. When the manifold supports a parallel spinor - Ricci-flatness is a required integrability condition for parallel spinors, the trace of Dirac Laplacian heat kernel of the spinor bundle can be recovered. For 4-manifolds, this corresponds to the spectral action, and realizes Einstein-Hilbert action as a stochastic flow.2024-01-01T03:08:36ZSita GakkharMatilde Marcollihttp://arxiv.org/abs/2401.00666v1On the $δ$-chromatic numbers of the Cartesian products of graphs2024-01-01T04:55:47ZIn this work, we study the $δ$-chromatic number of a graph which is the chromatic number of the $δ$-complement of a graph. We give a structure of the $δ$-complements and sharp bounds on the $δ$-chromatic numbers of the Cartesian products of graphs. Furthermore, we compute the $δ$-chromatic numbers of various classes of Cartesian product graphs, including the Cartesian products between cycles, paths, and stars.2024-01-01T04:55:47ZWipawee TangjaiWitsarut Pho-onPanupong Vichitkunakornhttp://arxiv.org/abs/2401.00669v5Analysis of the electromagnetic form factors and the radiative decays of the vector heavy-light mesons2024-03-04T02:55:30ZIn this article, we analyze the electromagnetic form factors of the vector heavy-light mesons to the pseudoscalar heavy-light mesons in the framework of three-point QCD sum rules, where the contributions of vacuum condensate terms $\langle\overline{q}q\rangle$, $\langle\overline{q}g_{s}σGq\rangle$, $\langle g_{s}^{2}G^{2}\rangle$, $\langle f^{3}G^{3}\rangle$ and $\langle\overline{q}q\rangle\langle g_{s}^{2}G^{2}\rangle$ are considered. With these results, we also obtain the radiative decay widths of the vector heavy-light mesons and then compare our results with those of other collaboration's. The final results about the radiative decay widths are $Γ(D^{*0}\to D^{0}γ)=1.74^{+0.40}_{-0.37}$ keV, $Γ(D^{*+}\to D^{+}γ)=0.17^{+0.08}_{-0.07}$ keV, $Γ(D_{s}^{*}\to D_{s}γ)=0.029^{+0.009}_{-0.008}$ keV, $Γ(B^{*0}\to B^{0}γ)=0.018^{+0.006}_{-0.005}$ keV, $Γ(B^{*+}\to B^{+}γ)=0.015^{+0.007}_{-0.007}$ keV and $Γ(B^{*}_{s}\to B_{s}γ)=0.016^{+0.003}_{-0.005}$ keV.2024-01-01T05:19:54ZPhys. Lett. B 852 (2024) 138624Jie LuGuo-Liang YuZhi-Gang WangBin Wu10.1016/j.physletb.2024.138624http://arxiv.org/abs/2401.00676v1Digger: Detecting Copyright Content Mis-usage in Large Language Model Training2024-01-01T06:04:52ZPre-training, which utilizes extensive and varied datasets, is a critical factor in the success of Large Language Models (LLMs) across numerous applications. However, the detailed makeup of these datasets is often not disclosed, leading to concerns about data security and potential misuse. This is particularly relevant when copyrighted material, still under legal protection, is used inappropriately, either intentionally or unintentionally, infringing on the rights of the authors.
In this paper, we introduce a detailed framework designed to detect and assess the presence of content from potentially copyrighted books within the training datasets of LLMs. This framework also provides a confidence estimation for the likelihood of each content sample's inclusion. To validate our approach, we conduct a series of simulated experiments, the results of which affirm the framework's effectiveness in identifying and addressing instances of content misuse in LLM training processes. Furthermore, we investigate the presence of recognizable quotes from famous literary works within these datasets. The outcomes of our study have significant implications for ensuring the ethical use of copyrighted materials in the development of LLMs, highlighting the need for more transparent and responsible data management practices in this field.2024-01-01T06:04:52ZHaodong LiGelei DengYi LiuKailong WangYuekang LiTianwei ZhangYang LiuGuoai XuGuosheng XuHaoyu Wanghttp://arxiv.org/abs/2401.00677v1Linear subspaces of the intersection of two quadrics via Kuznetsov component2024-01-01T06:13:30ZLet $Q_i(i=1,2)$ be $2g$ dimensional quadrics in $\mathbb{P}^{2g+1}$ and let $Y$ be the smooth intersection $Q_1\cap Q_2$. We associate the linear subspace in $Y$ with vector bundles on the hyperelliptic curve $C$ of genus $g$ by the left adjoint functor of $Φ:D^b(C)\rightarrow D^b(Y)$. As an application, we give a different proof of the classification of line bundles and stable bundles of rank $2$ on hyperelliptic curves given by Desale and Ramanan. When $g=3$, we show that the projection functor induces a closed embedding $α:Y\rightarrow SU^s_C(4,h)$ into the moduli space of stable bundles on $C$ of rank $4$ of fixed determinant.2024-01-01T06:13:30Z15 pagesYanjie LiShizhuo Zhanghttp://arxiv.org/abs/2401.00681v2Mitigating Procrastination in Spatial Crowdsourcing Via Efficient Scheduling Algorithm2024-02-07T14:07:38ZSeveral works related to spatial crowdsourcing have been proposed in the direction where the task executers are to perform the tasks within the stipulated deadlines. Though the deadlines are set, it may be a practical scenario that majority of the task executers submit the tasks as late as possible. This situation where the task executers may delay their task submission is termed as procrastination in behavioural economics. In many applications, these late submission of tasks may be problematic for task providers. So here, the participating agents (both task providers and task executers) are articulated with the procrastination issue. In literature, how to prevent this procrastination within the deadline is not addressed in spatial crowdsourcing scenario. However, in a bipartite graph setting one procrastination aware scheduling is proposed but balanced job (task and job will synonymously be used) distribution in different slots (also termed as schedules) is not considered there. In this paper, a procrastination aware scheduling of jobs is proliferated by proposing an (randomized) algorithm in spatial crowdsourcing scenario. Our algorithm ensures that balancing of jobs in different schedules are maintained. Our scheme is compared with the existing algorithm through extensive simulation and in terms of balancing effect, our proposed algorithm outperforms the existing one. Analytically it is shown that our proposed algorithm maintains the balanced distribution.2024-01-01T06:41:20ZNaren DebnathSajal MukhopadhyayFatos Xhafahttp://arxiv.org/abs/2401.00682v1The Smooth Trajectory Estimator for LMB Filters2024-01-01T06:45:40ZThis paper proposes a smooth-trajectory estimator for the labelled multi-Bernoulli (LMB) filter by exploiting the special structure of the generalised labelled multi-Bernoulli (GLMB) filter. We devise a simple and intuitive approach to store the best association map when approximating the GLMB random finite set (RFS) to the LMB RFS. In particular, we construct a smooth-trajectory estimator (i.e., an estimator over the entire trajectories of labelled estimates) for the LMB filter based on the history of the best association map and all of the measurements up to the current time. Experimental results under two challenging scenarios demonstrate significant tracking accuracy improvements with negligible additional computational time compared to the conventional LMB filter. The source code is publicly available at https://tinyurl.com/ste-lmb, aimed at promoting advancements in MOT algorithms.2024-01-01T06:45:40Z6 pages, 5 figures. Presented at The 12th IEEE International Conference on Control, Automation and Information Sciences (ICCAIS 2023), Nov 2023, Hanoi, VietnamHoa Van NguyenTran Thien Dat NguyenChangbeom ShimMarzhar Anuar10.1109/ICCAIS59597.2023.10382267http://arxiv.org/abs/2401.00689v1Large language model for Bible sentiment analysis: Sermon on the Mount2024-01-01T07:35:29ZThe revolution of natural language processing via large language models has motivated its use in multidisciplinary areas that include social sciences and humanities and more specifically, comparative religion. Sentiment analysis provides a mechanism to study the emotions expressed in text. Recently, sentiment analysis has been used to study and compare translations of the Bhagavad Gita, which is a fundamental and sacred Hindu text. In this study, we use sentiment analysis for studying selected chapters of the Bible. These chapters are known as the Sermon on the Mount. We utilize a pre-trained language model for sentiment analysis by reviewing five translations of the Sermon on the Mount, which include the King James version, the New International Version, the New Revised Standard Version, the Lamsa Version, and the Basic English Version. We provide a chapter-by-chapter and verse-by-verse comparison using sentiment and semantic analysis and review the major sentiments expressed. Our results highlight the varying sentiments across the chapters and verses. We found that the vocabulary of the respective translations is significantly different. We detected different levels of humour, optimism, and empathy in the respective chapters that were used by Jesus to deliver his message.2024-01-01T07:35:29ZMahek VoraTom BlauVansh KachhwalAshu M. G. SoloRohitash Chandrahttp://arxiv.org/abs/2401.00696v2$Z^\prime$ induced forward dominant processes in $μ$TRISTAN experiment2024-04-16T04:19:49ZGeneral $U(1)$ extension of the Standard Model (SM) is a well motivated beyond the Standard Model(BSM) scenario where three generations of right handed neutrinos (RHNs) are introduced to cancel gauge and mixed gauge-gravity anomalies. After the $U(1)_X$ is broken, RHNs participate in the seesaw mechanism to generate light neutrino masses satisfying neutrino oscillation data. In addition to that, a neutral gauge boson $Z^\prime$ is evolved which interacts with the left and right handed fermions differently manifesting chiral nature of the model which could be probed in future collider experiments. As a result, if we consider $μ^+ e^-$ and $μ^+ μ^+$ collisions in $μ$TRISTAN experiment $Z^\prime$ mediated $2\to2$ scattering will appear in $t-$ and $u-$channels depending on the initial and final states being accompanied by the photon and $Z$ mediated interactions. This will result well motivated resulting forward dominant scenarios giving rise to sizable left-right asymmetry. Estimating constrains on general $U(1)$ coupling from LEP-II and LHC for different $U(1)_X$ charges, we calculate differential and integrated scattering cross section and left-right asymmetry for $μ^+ e^- \to μ^+ e^-$ and $μ^+ μ^+ \to μ^+ μ^+$ processes which could be probed at $μ$TRISTAN experiment further enlightening the interaction between $Z^\prime$ and charged leptons and the $U(1)_X$ breaking scale.2024-01-01T08:19:21Z18 pagesPhys.Lett.B 851 (2024) 138577Arindam DasYuta Orikasa10.1016/j.physletb.2024.138577http://arxiv.org/abs/2401.00636v2Periodic and quasi-motivic pencils of flat connections2024-05-04T13:02:42ZWe introduce a new notion of a periodic pencil of flat connections on a smooth algebraic variety $X$. This is a family $\nabla(s_1,...,s_n)$ of flat connections on a trivial vector bundle on $X$ depending linearly on parameters $s_1,...,s_n$ and generically invariant, up to isomorphism, under the shifts $s_i\mapsto s_i+1$ for all $i$. If in addition $\nabla$ has regular singularities, we call it a quasi-motivic pencil. We use tools from complex analysis to establish various remarkable properties of such pencils over $\mathbb C$. For example, we show that the monodromy of a quasi-motivic pencil is defined over the field of algebraic functions in $e^{2πis_j}$, and that its singularities are constrained to an arrangement of hyperplanes with integer normal vectors. Then we show that many important examples of families of flat connections, such as Knizhnik-Zamolodchikov, Dunkl, and Casimir connections, are quasi-motivic and thus periodic pencils.
Besides being interesting in its own right, the periodic property of a pencil of flat connections turns out to be very useful in computing the eigenvalues of the $p$-curvature of its reduction to positive characteristic. This will be done in our forthcoming paper.2024-01-01T02:16:26Z29 pages, latex; in v2 small corrections and new Section 4.8Pavel EtingofAlexander Varchenkohttp://arxiv.org/abs/2401.00631v2Edge AI as a Service with Coordinated Deep Neural Networks2024-08-21T17:47:53ZAs artificial intelligence (AI) applications continue to expand in next-generation networks, there is a growing need for deep neural network (DNN) models. Although DNN models deployed at the edge are promising for providing AI as a service with low latency, their cooperation is yet to be explored. In this paper, we consider that DNN service providers share their computing resources as well as their models' parameters and allow other DNNs to offload their computations without mirroring. We propose a novel algorithm called coordinated DNNs on edge (\textbf{CoDE}) that facilitates coordination among DNN services by establishing new inference paths. CoDE aims to find the optimal path, which is the path with the highest possible reward, by creating multi-task DNNs from individual models. The reward reflects the inference throughput and model accuracy. With CoDE, DNN models can make new paths for inference by using their own or other models' parameters. We then evaluate the performance of CoDE through numerical experiments. The results demonstrate a $40\%$ increase in the inference throughput while degrading the average accuracy by only $2.3\%$. Experiments show that CoDE enhances the inference throughput and, achieves higher precision compared to a state-of-the-art existing method.2024-01-01T01:54:53ZAlireza MalekiHamed Shah-MansouriBabak H. Khalajhttp://arxiv.org/abs/2401.00699v3Quantum walk on simplicial complexes for simplicial community detection2024-04-26T14:25:20ZQuantum walks have emerged as a transformative paradigm in quantum information processing and can be applied to various graph problems. This study explores discrete-time quantum walks on simplicial complexes, a higher-order generalization of graph structures. Simplicial complexes, encoding higher-order interactions through simplices, offer a richer topological representation of complex systems. Since the conventional classical random walk cannot directly detect community structures, we present a quantum walk algorithm to detect higher-order community structures called simplicial communities. We utilize the Fourier coin to produce entangled translation states among adjacent simplices in a simplicial complex. The potential of our quantum algorithm is tested on Zachary's karate club network. This study may contribute to understanding complex systems at the intersection of algebraic topology and quantum walk algorithms.2024-01-01T08:43:43Z14 pages, manuscript revisedQuantum Information Processing 23, 199 (2024)Euijun Song10.1007/s11128-024-04415-9http://arxiv.org/abs/2401.00624v3Semi-Confirmatory Factor Analysis for High-Dimensional Data with Interconnected Community Structures2024-10-06T17:54:25ZConfirmatory factor analysis (CFA) is a statistical method for identifying and confirming the presence of latent factors among observed variables through the analysis of their covariance structure. Compared to alternative factor models, CFA offers interpretable common factors with enhanced specificity and a more adaptable approach to covariance structure modeling. However, the application of CFA has been limited by the requirement for prior knowledge about "non-zero loadings" and by the lack of computational scalability (e.g., it can be computationally intractable for hundreds of observed variables). We propose a data-driven semi-confirmatory factor analysis (SCFA) model that attempts to alleviate these limitations. SCFA automatically specifies "non-zero loadings" by learning the network structure of the large covariance matrix of observed variables, and then offers closed-form estimators for factor loadings, factor scores, covariances between common factors, and variances between errors using the likelihood method. Therefore, SCFA is applicable to high-throughput datasets (e.g., hundreds of thousands of observed variables) without requiring prior knowledge about "non-zero loadings". Through an extensive simulation analysis benchmarking against standard packages, SCFA exhibits superior performance in estimating model parameters with a much-reduced computational time. We illustrate its practical application through factor analysis on two high-dimensional RNA-seq gene expression datasets.2024-01-01T00:56:54ZYifan YangTianzhou MaChuan BiShuo Chenhttp://arxiv.org/abs/2401.00675v2Exotic synchronization in continuous time crystals outside the symmetric subspace2024-11-27T08:43:20ZExploring continuous time crystals (CTCs) within the symmetric subspace of spin systems has been a subject of intensive research in recent times. Thus far, the stability of the time-crystal phase outside the symmetric subspace in such spin systems has gone largely unexplored. Here, we investigate the effect of including the asymmetric subspaces on the dynamics of CTCs in a driven dissipative spin model. This results in multistability, and the dynamics becomes dependent on the initial state. Remarkably, this multistability leads to exotic synchronization regimes such as chimera states and cluster synchronization in an ensemble of coupled identical CTCs. Interestingly, it leads to other nonlinear phenomena such as oscillation death and signature of chaos.2024-01-01T06:00:51Z11 pages, close to accepted versionPhys. Rev. Lett. 133, 260403 (2024)Parvinder SolankiMidhun KrishnaMichal HajdušekChristoph BruderSai Vinjanampathy10.1103/PhysRevLett.133.260403http://arxiv.org/abs/2401.00672v4OBK-RCM: Accelerated Orthogonal Block Kaczmarz Algorithm via RCM Reordering and Dynamic Grouping for Sparse Linear Systems2025-06-21T14:53:08ZExisting block Kaczmarz methods face challenges in balancing computational efficiency and convergence for large sparse linear systems with scattered nonzero patterns, due to costly partitioning strategies and non-orthogonal projections. In this paper, we propose the orthogonal block Kaczmarz (OBK-RCM) algorithm with the Reverse Cuthill-McKee (RCM), which integrates the RCM reordering with a novel orthogonal block partitioning strategy. RCM transforms sparse matrices into banded structures to enhance inter-block orthogonality, while dynamic grouping of mutually orthogonal blocks based on angle cosine thresholds reduces iterative complexity. In addition, two extended versions (SOBK-RCM and UOBK-RCM) are proposed to deal with non-square systems by constructing extended matrices without sacrificing sparsity. This work offers a practical framework for efficient sparse linear algebra solvers. Experiments on 33 real-world and synthetic matrices show that OBK-RCM achieves 10-50 times faster CPU time (up to several hundred) and 50-90% fewer iterations than state-of-the-art methods (RBK,RBK(k),GREBK(k),aRBK), especially for scattered sparse structures in most cases. Theoretical analysis confirms linear convergence, driven by hyperplane orthogonality.2024-01-01T05:31:18Z29 pages, 12 figures, 10 tablesYu-Fang LiangHou-Biao Li