HACKOBAR_ABOUT

AI News Archive

183 items · 50 per page
1OPSRD Enables On-Policy Self-Distillation Using Expert Role Prompting10m ago0.452ArchitectureIQ Benchmark Measures LLM Training Intuition vs Human Experts10m ago0.283SpeechConversationBench Evaluates Multi-Turn Reasoning in Speech-to-Speech Models10m ago0.404OSWorld-Science Benchmark for Computer-Use Agents in Scientific Workflows10m ago0.36
5
BiFE for Efficient CPU-Only Branching Policies via LLMs
41m ago
0.18
6Framework for Distilling Knowledge in Complex Agentic Systems41m ago0.21
7Spectral Optimization for Controlling Safety Instruction Strength41m ago0.24
8FedLAFP Framework for Personalized Federated Fine-Tuning41m ago0.27
9Last-Chance Policy Identification for Agents Under Resource Depletion41m ago0.21
10OpenAI Agents Escaped Sandboxes via Zero-Day Vulnerabilities41m ago0.35
11SAKI Method Optimizes On-Policy Distillation via Maximal-Coupling-Routed Supervision1h ago0.24
12ARCagent Calibrates Clinical RAG for Conflicting Medical Guidelines1h ago0.27
13Self-Correction Trade-offs in 29 Open-Weight LLMs1h ago0.31
14Attuner Enables Recomputation-Free KV Cache Reuse via Query-Side Adaptation1h ago0.34
15JEV Model Cultural Alignment Shifts with Persona and Language1h ago0.21
16Counterfactual Diagnostics Separate Visual Sensitivity from Claim Persistence in LVLMs1h ago0.24
17HyperZip Uses Diffusion LLMs and Hypernetworks for Faster Compression1h ago0.21
18Identifying Attention Heads Responsible for LLM Sycophancy1h ago0.27
19ContextProgress-Bench Evaluates Progress Reward Models in Long-Horizon Tasks1h ago0.18
20LLM-Guided Pruning Fixes Geometry-Semantic Mismatch in ANN Graphs1h ago0.21
21Factor Analysis of 1,618 Models Reveals Partial Intelligence Interpretability2h ago0.18
22HeurEvo Automates Hybrid Solver-Augmented Heuristics via Co-Evolution2h ago0.21
23LUDI Framework Scales Diffusion Language Models with Per-Token Embeddings2h ago0.35
24RankBuffer Optimizes Open-Ended Generation via Reusable Quality Buffers2h ago0.27
25StateTape Rewrites Coding Agent Context via Repository State Changes2h ago0.31
26FinRT Framework Automates Adversarial Prompt Generation for Finance Models2h ago0.31
27CoRe Framework Mitigates Latent Reward Hacking in Video Diffusion2h ago0.28
28evalstats Tooling Provides Calibrated Statistical Inference for LLM Judges2h ago0.24
29TrustSwap Reveals Source-Trust Shortcuts in Fact-Checking RL Agents2h ago0.28
30OLIVE Distillation Method Improves Reasoning Performance via Teacher Continuations2h ago0.35
31MemFold Optimizes Fixed-Budget Soft Memory for Long-Context Personalization3h ago0.31
32LLMs Suffer Reliability Issues with Non-Canonical Multi-Valued Relations3h ago0.21
33BreakingWeb Benchmarks Browser-Use Agents via Controlled Environment Interventions3h ago0.24
34AdaLCPI Attack Reconstructs Indirect Prompt Injections from Fragments3h ago0.35
35CineSubBench Evaluates Long-Context Multilingual Film Understanding3h ago0.18
36Framework Desktop Preorders Open for AMD Ryzen AI Max 400 Series3h ago0.06
37Dense Retrievers Exhibit Political and Dialect Bias in Query Responses3h ago0.35
38PhenoAIR Uses Multi-Agent Reasoning for Cell Painting MOA Prediction3h ago0.24
39FigAct Transforms Static Scientific Figures Into Interactive Visual Presentations3h ago0.31
40GPT Models Generate More Edit-Inducing Questions Than Human Reviewers3h ago0.18
41ATRI Framework Mitigates LLM Self-Evolution Degeneration via Information Gain3h ago0.28
42Sequential Interfaces and Reasoning Scaffolds Improve LLM Agent Decision-Making4h ago0.28
43Character-Level LSTM Models for Sinhala Sandhi Splitting4h ago0.18
44Lookahead-R Optimizes Tool Retrieval Using Execution-Aware Surrogate World Models4h ago0.35
45Training-Free Human Activity Recognition Fails to Match Supervised Performance4h ago0.21
46Automated Preoperative Screw Planning for Pelvic Fractures4h ago0.18
47ThuRunel Decouples Elicitation and Judgment in Advisory Agents4h ago0.28
48MemEvo Automates Streaming Video Memory Mechanism Discovery4h ago0.21
49Reliable Parallel Decoding for Masked Diffusion Language Models4h ago0.25
50Omni-Decision uses Evidence-Ledger Planning for Multimodal Agents4h ago0.54
page 1 / 4next →