<?xml version="1.0" encoding="UTF-8"?><feed xmlns="http://www.w3.org/2005/Atom"><title>TechAIOrbit</title><id>https://techaiorbit.com/</id><link href="https://techaiorbit.com/"/><link rel="self" href="https://techaiorbit.com/atom.xml"/><updated>2026-09-09T18:21:40.000Z</updated><entry><title>Apple Watch Ultra 4</title><id>https://techaiorbit.com/news/apple-watch-ultra-4-49631121/</id><link href="https://techaiorbit.com/news/apple-watch-ultra-4-49631121/"/><link rel="related" href="https://www.apple.com/apple-watch-ultra-4/"/><updated>2026-09-09T18:21:40.000Z</updated><summary>Source-linked news record from Hacker News.</summary></entry><entry><title>Apple Unveils iPhone Duo</title><id>https://techaiorbit.com/news/apple-unveils-iphone-duo-49630964/</id><link href="https://techaiorbit.com/news/apple-unveils-iphone-duo-49630964/"/><link rel="related" href="https://www.apple.com/newsroom/2026/09/apple-unveils-iphone-duo/"/><updated>2026-09-09T18:17:06.000Z</updated><summary>Source-linked news record from Hacker News.</summary></entry><entry><title>iPhone Duo</title><id>https://techaiorbit.com/news/iphone-duo-49630931/</id><link href="https://techaiorbit.com/news/iphone-duo-49630931/"/><link rel="related" href="https://www.apple.com/iphone-duo/"/><updated>2026-09-09T18:15:43.000Z</updated><summary>Source-linked news record from Hacker News.</summary></entry><entry><title>Apple Introduces AirPods 5</title><id>https://techaiorbit.com/news/apple-introduces-airpods-5-49630253/</id><link href="https://techaiorbit.com/news/apple-introduces-airpods-5-49630253/"/><link rel="related" href="https://www.apple.com/newsroom/2026/09/apple-introduces-airpods-5-with-best-in-class-open-ear-active-noise-cancellation/"/><updated>2026-09-09T17:39:24.000Z</updated><summary>Source-linked news record from Hacker News.</summary></entry><entry><title>Qwen 3.8 follows GPT-5.5 Pro reasoning prefills</title><id>https://techaiorbit.com/news/qwen-38-follows-gpt-55-pro-reasoning-prefills-49630026/</id><link href="https://techaiorbit.com/news/qwen-38-follows-gpt-55-pro-reasoning-prefills-49630026/"/><link rel="related" href="https://gist.github.com/wsxiaoys/e0286dc6bb624ff5fdf49e7f4c528ba3"/><updated>2026-09-09T17:24:28.000Z</updated><summary>Source-linked news record from Hacker News.</summary></entry><entry><title>Procedural Graphs: Self-Evolving Execution Structures for LLM Agents</title><id>https://techaiorbit.com/news/procedural-graphs-self-evolving-execution-structures-for-llm-agents-49629868/</id><link href="https://techaiorbit.com/news/procedural-graphs-self-evolving-execution-structures-for-llm-agents-49629868/"/><link rel="related" href="https://academy.dair.ai/papers/procedural-graphs-self-evolving-execution-structures-for-llm-agents-2609.09153"/><updated>2026-09-09T17:13:52.000Z</updated><summary>Source-linked news record from Hacker News.</summary></entry><entry><title>Why Emacs Consult async searches feel slow and how to speed them up</title><id>https://techaiorbit.com/news/why-emacs-consult-async-searches-feel-slow-and-how-to-speed-them-up-49629865/</id><link href="https://techaiorbit.com/news/why-emacs-consult-async-searches-feel-slow-and-how-to-speed-them-up-49629865/"/><link rel="related" href="https://www.jamescherti.com/emacs-consult-speed-async-searche-grep-ripgrep-fd-find/"/><updated>2026-09-09T17:13:43.000Z</updated><summary>Source-linked news record from Hacker News.</summary></entry><entry><title>Understanding the recent DDoS attack against Read the Docs</title><id>https://techaiorbit.com/news/understanding-the-recent-ddos-attack-against-read-the-docs-49628614/</id><link href="https://techaiorbit.com/news/understanding-the-recent-ddos-attack-against-read-the-docs-49628614/"/><link rel="related" href="https://about.readthedocs.com/blog/2026/09/2026-ddos-attack/"/><updated>2026-09-09T15:55:53.000Z</updated><summary>Source-linked news record from Hacker News.</summary></entry><entry><title>GNU Radio in the browser</title><id>https://techaiorbit.com/news/gnu-radio-in-the-browser-49628576/</id><link href="https://techaiorbit.com/news/gnu-radio-in-the-browser-49628576/"/><link rel="related" href="https://gnuradioworld.com/"/><updated>2026-09-09T15:53:06.000Z</updated><summary>Source-linked news record from Hacker News.</summary></entry><entry><title>Planet Labs&#39; open satellite feed</title><id>https://techaiorbit.com/news/planet-labs-open-satellite-feed-49628429/</id><link href="https://techaiorbit.com/news/planet-labs-open-satellite-feed-49628429/"/><link rel="related" href="https://tech.marksblogg.com/planet-labs-open-satellite-feed.html"/><updated>2026-09-09T15:44:09.000Z</updated><summary>Source-linked news record from Hacker News.</summary></entry><entry><title>We accidentally built a synthetic cell factory</title><id>https://techaiorbit.com/news/we-accidentally-built-a-synthetic-cell-factory-49628290/</id><link href="https://techaiorbit.com/news/we-accidentally-built-a-synthetic-cell-factory-49628290/"/><link rel="related" href="https://bnext.bio/post/we-accidentally-built-a-synthetic-cell-factory"/><updated>2026-09-09T15:34:51.000Z</updated><summary>Source-linked news record from Hacker News.</summary></entry><entry><title>GPT-6 Astra, looped transformers, and hidden reasoning</title><id>https://techaiorbit.com/news/gpt-6-astra-looped-transformers-and-hidden-reasoning-49627370/</id><link href="https://techaiorbit.com/news/gpt-6-astra-looped-transformers-and-hidden-reasoning-49627370/"/><link rel="related" href="https://magazine.sebastianraschka.com/p/gpt-6-astra-looped-transformers-and"/><updated>2026-09-09T14:37:47.000Z</updated><summary>Source-linked news record from Hacker News.</summary></entry><entry><title>Tailwind Labs is joining Shopify</title><id>https://techaiorbit.com/news/tailwind-labs-is-joining-shopify-49626190/</id><link href="https://techaiorbit.com/news/tailwind-labs-is-joining-shopify-49626190/"/><link rel="related" href="https://tailwindcss.com/blog/tailwind-is-joining-shopify"/><updated>2026-09-09T13:27:11.000Z</updated><summary>Source-linked news record from Hacker News.</summary></entry><entry><title>How I advertise malicious software on Google Ads</title><id>https://techaiorbit.com/news/how-i-advertise-malicious-software-on-google-ads-49624856/</id><link href="https://techaiorbit.com/news/how-i-advertise-malicious-software-on-google-ads-49624856/"/><link rel="related" href="https://xlii.space/eng/malicious-software-on-google-ads/"/><updated>2026-09-09T11:43:21.000Z</updated><summary>Source-linked news record from Hacker News.</summary></entry><entry><title>Desert Ant Labs: local, fast models that run on device</title><id>https://techaiorbit.com/news/desert-ant-labs-local-fast-models-that-run-on-device-49624823/</id><link href="https://techaiorbit.com/news/desert-ant-labs-local-fast-models-that-run-on-device-49624823/"/><link rel="related" href="https://desertant.com/blog/introducing-desert-ant-labs/"/><updated>2026-09-09T11:39:46.000Z</updated><summary>Source-linked news record from Hacker News.</summary></entry><entry><title>Generating the P3 Tiling</title><id>https://techaiorbit.com/news/generating-the-p3-tiling-49614575/</id><link href="https://techaiorbit.com/news/generating-the-p3-tiling-49614575/"/><link rel="related" href="https://k-monk.org/blog/generating-the-p3-tiling/"/><updated>2026-09-08T18:27:32.000Z</updated><summary>Source-linked news record from Hacker News.</summary></entry><entry><title>What do Visa and Mastercard do? An intro to card networks</title><id>https://techaiorbit.com/news/what-do-visa-and-mastercard-do-an-intro-to-card-networks-49614280/</id><link href="https://techaiorbit.com/news/what-do-visa-and-mastercard-do-an-intro-to-card-networks-49614280/"/><link rel="related" href="https://tautology.town/2026/06/01/card-networks.html"/><updated>2026-09-08T18:11:05.000Z</updated><summary>Source-linked news record from Hacker News.</summary></entry><entry><title>San Francisco institution &#39;heartbroken and furious&#39; after erasure of public art</title><id>https://techaiorbit.com/news/san-francisco-institution-heartbroken-and-furious-after-erasure-of-public-art-49600627/</id><link href="https://techaiorbit.com/news/san-francisco-institution-heartbroken-and-furious-after-erasure-of-public-art-49600627/"/><link rel="related" href="https://www.sfgate.com/local/article/clarion-alley-murals-erased-22420365.php"/><updated>2026-09-07T17:19:14.000Z</updated><summary>Source-linked news record from Hacker News.</summary></entry><entry><title>Bespoke: A programming language for people who say please</title><id>https://techaiorbit.com/news/bespoke-a-programming-language-for-people-who-say-please-49584361/</id><link href="https://techaiorbit.com/news/bespoke-a-programming-language-for-people-who-say-please-49584361/"/><link rel="related" href="https://blog.hofstede.it/bespoke-a-programming-language-for-people-who-say-please/"/><updated>2026-09-06T08:13:53.000Z</updated><summary>Source-linked news record from Hacker News.</summary></entry><entry><title>Neural document expansion for ad-hoc information retrieval</title><id>https://techaiorbit.com/research/neural-document-expansion-for-ad-hoc-information-retrieval-201214005/</id><link href="https://techaiorbit.com/research/neural-document-expansion-for-ad-hoc-information-retrieval-201214005/"/><link rel="related" href="https://arxiv.org/abs/2012.14005"/><updated>2020-12-27T20:00:08.000Z</updated><summary>Recently, Nogueira et al. [2019] proposed a new approach to document expansion based on a neural Seq2Seq model, showing significant improvement on short text retrieval task. However, this approach needs a large amount of in-domain training data. In this paper, we show that this neural document expansion approach can be effectively adapted to standard IR tasks, where labels are scarce and many long documents are present.</summary></entry><entry><title>I like fish, especially dolphins: Addressing Contradictions in Dialogue Modeling</title><id>https://techaiorbit.com/research/i-like-fish-especially-dolphins-addressing-contradictions-in-dialogue-modeling-201213391/</id><link href="https://techaiorbit.com/research/i-like-fish-especially-dolphins-addressing-contradictions-in-dialogue-modeling-201213391/"/><link rel="related" href="https://arxiv.org/abs/2012.13391"/><updated>2020-12-24T18:47:49.000Z</updated><summary>To quantify how well natural language understanding models can capture consistency in a general conversation, we introduce the DialoguE COntradiction DEtection task (DECODE) and a new conversational dataset containing both human-human and human-bot contradictory dialogues. We then compare a structured utterance-based approach of using pre-trained Transformer models for contradiction detection with the typical unstructured approach. Results reveal that: (i) our newly collected dataset is notably more effective at providing supervision for the dialogue contradiction detection task than existing NLI data including those aimed to cover the dialogue domain; (ii) the structured utterance-based approach is more robust and transferable on both analysis and out-of-distribution dialogues than its unstructured counterpart. We also show that our best contradiction detection model correlates well with human judgments and further provide evidence for its usage in both automatically evaluating and improving the consistency of state-of-the-art generative chatbots.</summary></entry><entry><title>Detecting Insincere Questions from Text: A Transfer Learning Approach</title><id>https://techaiorbit.com/research/detecting-insincere-questions-from-text-a-transfer-learning-approach-201207587/</id><link href="https://techaiorbit.com/research/detecting-insincere-questions-from-text-a-transfer-learning-approach-201207587/"/><link rel="related" href="https://arxiv.org/abs/2012.07587"/><updated>2020-12-07T15:03:48.000Z</updated><summary>The internet today has become an unrivalled source of information where people converse on content based websites such as Quora, Reddit, StackOverflow and Twitter asking doubts and sharing knowledge with the world. A major arising problem with such websites is the proliferation of toxic comments or instances of insincerity wherein the users instead of maintaining a sincere motive indulge in spreading toxic and divisive content. The straightforward course of action in confronting this situation is detecting such content beforehand and preventing it from subsisting online. In recent times Transfer Learning in Natural Language Processing has seen an unprecedented growth. Today with the existence of transformers and various state of the art innovations, a tremendous growth has been made in various NLP domains. The introduction of BERT has caused quite a stir in the NLP community. As mentioned, when published, BERT dominated performance benchmarks and thereby inspired many other authors to experiment with it and publish similar models. This led to the development of a whole BERT-family, each member being specialized on a different task. In this paper we solve the Insincere Questions Classification problem by fine tuning four cutting age models viz BERT, RoBERTa, DistilBERT and ALBERT.</summary></entry><entry><title>Answer Span Correction in Machine Reading Comprehension</title><id>https://techaiorbit.com/research/answer-span-correction-in-machine-reading-comprehension-201103435/</id><link href="https://techaiorbit.com/research/answer-span-correction-in-machine-reading-comprehension-201103435/"/><link rel="related" href="https://arxiv.org/abs/2011.03435"/><updated>2020-11-06T15:31:07.000Z</updated><summary>Answer validation in machine reading comprehension (MRC) consists of verifying an extracted answer against an input context and question pair. Previous work has looked at re-assessing the &quot;answerability&quot; of the question given the extracted answer. Here we address a different problem: the tendency of existing MRC systems to produce partially correct answers when presented with answerable questions. We explore the nature of such errors and propose a post-processing correction method that yields statistically significant performance improvements over state-of-the-art MRC systems in both monolingual and multilingual evaluation.</summary></entry><entry><title>A Cross-lingual Natural Language Processing Framework for Infodemic Management</title><id>https://techaiorbit.com/research/a-cross-lingual-natural-language-processing-framework-for-infodemic-management-201016357/</id><link href="https://techaiorbit.com/research/a-cross-lingual-natural-language-processing-framework-for-infodemic-management-201016357/"/><link rel="related" href="https://arxiv.org/abs/2010.16357"/><updated>2020-10-30T16:26:35.000Z</updated><summary>The COVID-19 pandemic has put immense pressure on health systems which are further strained due to the misinformation surrounding it. Under such a situation, providing the right information at the right time is crucial. There is a growing demand for the management of information spread using Artificial Intelligence. Hence, we have exploited the potential of Natural Language Processing for identifying relevant information that needs to be disseminated amongst the masses. In this work, we present a novel Cross-lingual Natural Language Processing framework to provide relevant information by matching daily news with trusted guidelines from the World Health Organization. The proposed pipeline deploys various techniques of NLP such as summarizers, word embeddings, and similarity metrics to provide users with news articles along with a corresponding healthcare guideline. A total of 36 models were evaluated and a combination of LexRank based summarizer on Word2Vec embedding with Word Mover distance metric outperformed all other models. This novel open-source approach can be used as a template for proactive dissemination of relevant healthcare information in the midst of misinformation spread associated with epidemics.</summary></entry><entry><title>NU-GAN: High resolution neural upsampling with GAN</title><id>https://techaiorbit.com/research/nu-gan-high-resolution-neural-upsampling-with-gan-201011362/</id><link href="https://techaiorbit.com/research/nu-gan-high-resolution-neural-upsampling-with-gan-201011362/"/><link rel="related" href="https://arxiv.org/abs/2010.11362"/><updated>2020-10-22T01:00:23.000Z</updated><summary>In this paper, we propose NU-GAN, a new method for resampling audio from lower to higher sampling rates (upsampling). Audio upsampling is an important problem since productionizing generative speech technology requires operating at high sampling rates. Such applications use audio at a resolution of 44.1 kHz or 48 kHz, whereas current speech synthesis methods are equipped to handle a maximum of 24 kHz resolution. NU-GAN takes a leap towards solving audio upsampling as a separate component in the text-to-speech (TTS) pipeline by leveraging techniques for audio generation using GANs. ABX preference tests indicate that our NU-GAN resampler is capable of resampling 22 kHz to 44.1 kHz audio that is distinguishable from original audio only 7.4% higher than random chance for single speaker dataset, and 10.8% higher than chance for multi-speaker dataset.</summary></entry><entry><title>Using Type Information to Improve Entity Coreference Resolution</title><id>https://techaiorbit.com/research/using-type-information-to-improve-entity-coreference-resolution-201005738/</id><link href="https://techaiorbit.com/research/using-type-information-to-improve-entity-coreference-resolution-201005738/"/><link rel="related" href="https://arxiv.org/abs/2010.05738"/><updated>2020-10-12T14:32:39.000Z</updated><summary>Coreference resolution (CR) is an essential part of discourse analysis. Most recently, neural approaches have been proposed to improve over SOTA models from earlier paradigms. So far none of the published neural models leverage external semantic knowledge such as type information. This paper offers the first such model and evaluation, demonstrating modest gains in accuracy by introducing either gold standard or predicted types. In the proposed approach, type information serves both to (1) improve mention representation and (2) create a soft type consistency check between coreference candidate mentions. Our evaluation covers two different grain sizes of types over four different benchmark corpora.</summary></entry><entry><title>Evaluating and Characterizing Human Rationales</title><id>https://techaiorbit.com/research/evaluating-and-characterizing-human-rationales-201004736/</id><link href="https://techaiorbit.com/research/evaluating-and-characterizing-human-rationales-201004736/"/><link rel="related" href="https://arxiv.org/abs/2010.04736"/><updated>2020-10-09T18:00:04.000Z</updated><summary>Two main approaches for evaluating the quality of machine-generated rationales are: 1) using human rationales as a gold standard; and 2) automated metrics based on how rationales affect model behavior. An open question, however, is how human rationales fare with these automatic metrics. Analyzing a variety of datasets and models, we find that human rationales do not necessarily perform well on these metrics. To unpack this finding, we propose improved metrics to account for model-dependent baseline performance. We then propose two methods to further characterize rationale quality, one based on model retraining and one on using &quot;fidelity curves&quot; to reveal properties such as irrelevance and redundancy. Our work leads to actionable suggestions for evaluating and characterizing rationales.</summary></entry><entry><title>Robustness and Reliability of Gender Bias Assessment in Word Embeddings: The Role of Base Pairs</title><id>https://techaiorbit.com/research/robustness-and-reliability-of-gender-bias-assessment-in-word-embeddings-the-role-of-base-pairs-201002847/</id><link href="https://techaiorbit.com/research/robustness-and-reliability-of-gender-bias-assessment-in-word-embeddings-the-role-of-base-pairs-201002847/"/><link rel="related" href="https://arxiv.org/abs/2010.02847"/><updated>2020-10-06T16:09:05.000Z</updated><summary>It has been shown that word embeddings can exhibit gender bias, and various methods have been proposed to quantify this. However, the extent to which the methods are capturing social stereotypes inherited from the data has been debated. Bias is a complex concept and there exist multiple ways to define it. Previous work has leveraged gender word pairs to measure bias and extract biased analogies. We show that the reliance on these gendered pairs has strong limitations: bias measures based off of them are not robust and cannot identify common types of real-world bias, whilst analogies utilising them are unsuitable indicators of bias. In particular, the well-known analogy &quot;man is to computer-programmer as woman is to homemaker&quot; is due to word similarity rather than societal bias. This has important implications for work on measuring bias in embeddings and related work debiasing embeddings.</summary></entry><entry><title>A Spherical Hidden Markov Model for Semantics-Rich Human Mobility Modeling</title><id>https://techaiorbit.com/research/a-spherical-hidden-markov-model-for-semantics-rich-human-mobility-modeling-201001986/</id><link href="https://techaiorbit.com/research/a-spherical-hidden-markov-model-for-semantics-rich-human-mobility-modeling-201001986/"/><link rel="related" href="https://arxiv.org/abs/2010.01986"/><updated>2020-10-05T13:18:38.000Z</updated><summary>We study the problem of modeling human mobility from semantic trace data, wherein each GPS record in a trace is associated with a text message that describes the user&#39;s activity. Existing methods fall short in unveiling human movement regularities, because they either do not model the text data at all or suffer from text sparsity severely. We propose SHMM, a multi-modal spherical hidden Markov model for semantics-rich human mobility modeling. Under the hidden Markov assumption, SHMM models the generation process of a given trace by jointly considering the observed location, time, and text at each step of the trace. The distinguishing characteristic of SHMM is the text modeling part. We use fixed-size vector representations to encode the semantics of the text messages, and model the generation of the l2-normalized text embeddings on a unit sphere with the von Mises-Fisher (vMF) distribution. Compared with other alternatives like multi-variate Gaussian, our choice of the vMF distribution not only incurs much fewer parameters, but also better leverages the discriminative power of text embeddings in a directional metric space. The parameter inference for the vMF distribution is non-trivial since it involves functional inversion of ratios of Bessel functions. We theoretically prove that: 1) the classical Expectation-Maximization algorithm can work with vMF distributions; and 2) while closed-form solutions are hard to be obtained for the M-step, Newton&#39;s method is guaranteed to converge to the optimal solution with quadratic convergence rate. We have performed extensive experiments on both synthetic and real-life data. The results on synthetic data verify our theoretical analysis; while the results on real-life data demonstrate that SHMM learns meaningful semantics-rich mobility models, outperforms state-of-the-art mobility models for next location prediction, and incurs lower training cost.</summary></entry><entry><title>GenAug: Data Augmentation for Finetuning Text Generators</title><id>https://techaiorbit.com/research/genaug-data-augmentation-for-finetuning-text-generators-201001794/</id><link href="https://techaiorbit.com/research/genaug-data-augmentation-for-finetuning-text-generators-201001794/"/><link rel="related" href="https://arxiv.org/abs/2010.01794"/><updated>2020-10-05T05:46:39.000Z</updated><summary>In this paper, we investigate data augmentation for text generation, which we call GenAug. Text generation and language modeling are important tasks within natural language processing, and are especially challenging for low-data regimes. We propose and evaluate various augmentation methods, including some that incorporate external knowledge, for finetuning GPT-2 on a subset of Yelp Reviews. We also examine the relationship between the amount of augmentation and the quality of the generated text. We utilize several metrics that evaluate important aspects of the generated text including its diversity and fluency. Our experiments demonstrate that insertion of character-level synthetic noise and keyword replacement with hypernyms are effective augmentation methods, and that the quality of generations improves to a peak at approximately three times the amount of original data.</summary></entry><entry><title>WeChat Neural Machine Translation Systems for WMT20</title><id>https://techaiorbit.com/research/wechat-neural-machine-translation-systems-for-wmt20-201000247/</id><link href="https://techaiorbit.com/research/wechat-neural-machine-translation-systems-for-wmt20-201000247/"/><link rel="related" href="https://arxiv.org/abs/2010.00247"/><updated>2020-10-01T08:15:09.000Z</updated><summary>We participate in the WMT 2020 shared news translation task on Chinese to English. Our system is based on the Transformer (Vaswani et al., 2017a) with effective variants and the DTMT (Meng and Zhang, 2019) architecture. In our experiments, we employ data selection, several synthetic data generation approaches (i.e., back-translation, knowledge distillation, and iterative in-domain knowledge transfer), advanced finetuning approaches and self-bleu based model ensemble. Our constrained Chinese to English system achieves 36.9 case-sensitive BLEU score, which is the highest among all submissions.</summary></entry><entry><title>Mitigating Gender Bias for Neural Dialogue Generation with Adversarial Learning</title><id>https://techaiorbit.com/research/mitigating-gender-bias-for-neural-dialogue-generation-with-adversarial-learning-200913028/</id><link href="https://techaiorbit.com/research/mitigating-gender-bias-for-neural-dialogue-generation-with-adversarial-learning-200913028/"/><link rel="related" href="https://arxiv.org/abs/2009.13028"/><updated>2020-09-28T02:46:59.000Z</updated><summary>Dialogue systems play an increasingly important role in various aspects of our daily life. It is evident from recent research that dialogue systems trained on human conversation data are biased. In particular, they can produce responses that reflect people&#39;s gender prejudice. Many debiasing methods have been developed for various NLP tasks, such as word embedding. However, they are not directly applicable to dialogue systems because they are likely to force dialogue models to generate similar responses for different genders. This greatly degrades the diversity of the generated responses and immensely hurts the performance of the dialogue models. In this paper, we propose a novel adversarial learning framework Debiased-Chat to train dialogue models free from gender bias while keeping their performance. Extensive experiments on two real-world conversation datasets show that our framework significantly reduces gender bias in dialogue models while maintaining the response quality. The implementation of the proposed framework is released.</summary></entry><entry><title>On Data Augmentation for Extreme Multi-label Classification</title><id>https://techaiorbit.com/research/on-data-augmentation-for-extreme-multi-label-classification-200910778/</id><link href="https://techaiorbit.com/research/on-data-augmentation-for-extreme-multi-label-classification-200910778/"/><link rel="related" href="https://arxiv.org/abs/2009.10778"/><updated>2020-09-22T19:31:08.000Z</updated><summary>In this paper, we focus on data augmentation for the extreme multi-label classification (XMC) problem. One of the most challenging issues of XMC is the long tail label distribution where even strong models suffer from insufficient supervision. To mitigate such label bias, we propose a simple and effective augmentation framework and a new state-of-the-art classifier. Our augmentation framework takes advantage of the pre-trained GPT-2 model to generate label-invariant perturbations of the input texts to augment the existing training data. As a result, it present substantial improvements over baseline models. Our contributions are two-factored: (1) we introduce a new state-of-the-art classifier that uses label attention with RoBERTa and combine it with our augmentation framework for further improvement; (2) we present a broad study on how effective are different augmentation methods in the XMC task.</summary></entry><entry><title>Unsupervised Text Generation by Learning from Search</title><id>https://techaiorbit.com/research/unsupervised-text-generation-by-learning-from-search-200708557/</id><link href="https://techaiorbit.com/research/unsupervised-text-generation-by-learning-from-search-200708557/"/><link rel="related" href="https://arxiv.org/abs/2007.08557"/><updated>2020-07-09T04:34:48.000Z</updated><summary>In this work, we present TGLS, a novel framework to unsupervised Text Generation by Learning from Search. We start by applying a strong search algorithm (in particular, simulated annealing) towards a heuristically defined objective that (roughly) estimates the quality of sentences. Then, a conditional generative model learns from the search results, and meanwhile smooth out the noise of search. The alternation between search and learning can be repeated for performance bootstrapping. We demonstrate the effectiveness of TGLS on two real-world natural language generation tasks, paraphrase generation and text formalization. Our model significantly outperforms unsupervised baseline methods in both tasks. Especially, it achieves comparable performance with the state-of-the-art supervised methods in paraphrase generation.</summary></entry><entry><title>ClarQ: A large-scale and diverse dataset for Clarification Question Generation</title><id>https://techaiorbit.com/research/clarq-a-large-scale-and-diverse-dataset-for-clarification-question-generation-200605986/</id><link href="https://techaiorbit.com/research/clarq-a-large-scale-and-diverse-dataset-for-clarification-question-generation-200605986/"/><link rel="related" href="https://arxiv.org/abs/2006.05986"/><updated>2020-06-10T17:56:50.000Z</updated><summary>Question answering and conversational systems are often baffled and need help clarifying certain ambiguities. However, limitations of existing datasets hinder the development of large-scale models capable of generating and utilising clarification questions. In order to overcome these limitations, we devise a novel bootstrapping framework (based on self-supervision) that assists in the creation of a diverse, large-scale dataset of clarification questions based on post-comment tuples extracted from stackexchange. The framework utilises a neural network based architecture for classifying clarification questions. It is a two-step method where the first aims to increase the precision of the classifier and second aims to increase its recall. We quantitatively demonstrate the utility of the newly created dataset by applying it to the downstream task of question-answering. The final dataset, ClarQ, consists of ~2M examples distributed across 173 domains of stackexchange. We release this dataset in order to foster research into the field of clarification question generation with the larger goal of enhancing dialog and question answering systems.</summary></entry><entry><title>CycleGT: Unsupervised Graph-to-Text and Text-to-Graph Generation via Cycle Training</title><id>https://techaiorbit.com/research/cyclegt-unsupervised-graph-to-text-and-text-to-graph-generation-via-cycle-training-200604702/</id><link href="https://techaiorbit.com/research/cyclegt-unsupervised-graph-to-text-and-text-to-graph-generation-via-cycle-training-200604702/"/><link rel="related" href="https://arxiv.org/abs/2006.04702"/><updated>2020-06-08T15:59:00.000Z</updated><summary>Two important tasks at the intersection of knowledge graphs and natural language processing are graph-to-text (G2T) and text-to-graph (T2G) conversion. Due to the difficulty and high cost of data collection, the supervised data available in the two fields are usually on the magnitude of tens of thousands, for example, 18K in the WebNLG~2017 dataset after preprocessing, which is far fewer than the millions of data for other tasks such as machine translation. Consequently, deep learning models for G2T and T2G suffer largely from scarce training data. We present CycleGT, an unsupervised training method that can bootstrap from fully non-parallel graph and text data, and iteratively back translate between the two forms. Experiments on WebNLG datasets show that our unsupervised model trained on the same number of data achieves performance on par with several fully supervised models. Further experiments on the non-parallel GenWiki dataset verify that our method performs the best among unsupervised baselines. This validates our framework as an effective approach to overcome the data scarcity problem in the fields of G2T and T2G. Our code is available at https://github.com/QipengGuo/CycleGT.</summary></entry><entry><title>Probing Neural Dialog Models for Conversational Understanding</title><id>https://techaiorbit.com/research/probing-neural-dialog-models-for-conversational-understanding-200608331/</id><link href="https://techaiorbit.com/research/probing-neural-dialog-models-for-conversational-understanding-200608331/"/><link rel="related" href="https://arxiv.org/abs/2006.08331"/><updated>2020-06-07T17:32:00.000Z</updated><summary>The predominant approach to open-domain dialog generation relies on end-to-end training of neural models on chat datasets. However, this approach provides little insight as to what these models learn (or do not learn) about engaging in dialog. In this study, we analyze the internal representations learned by neural open-domain dialog systems and evaluate the quality of these representations for learning basic conversational skills. Our results suggest that standard open-domain dialog systems struggle with answering questions, inferring contradiction, and determining the topic of conversation, among other tasks. We also find that the dyadic, turn-taking nature of dialog is not fully leveraged by these models. By exploring these limitations, we highlight the need for additional research into architectures and training methods that can better capture high-level information about dialog.</summary></entry><entry><title>Conditioning LSTM Decoder and Bi-directional Attention Based Question Answering System</title><id>https://techaiorbit.com/research/conditioning-lstm-decoder-and-bi-directional-attention-based-question-answering-system-190502019/</id><link href="https://techaiorbit.com/research/conditioning-lstm-decoder-and-bi-directional-attention-based-question-answering-system-190502019/"/><link rel="related" href="https://arxiv.org/abs/1905.02019"/><updated>2019-05-02T01:07:20.000Z</updated><summary>Applying neural-networks on Question Answering has gained increasing popularity in recent years. In this paper, I implemented a model with Bi-directional attention flow layer, connected with a Multi-layer LSTM encoder, connected with one start-index decoder and one conditioning end-index decoder. I introduce a new end-index decoder layer, conditioning on start-index output. The Experiment shows this has increased model performance by 15.16%. For prediction, I proposed a new smart-span equation, rewarding both short answer length and high probability in start-index and end-index, which further improved the prediction accuracy. The best single model achieves an F1 score of 73.97% and EM score of 64.95% on test set.</summary></entry><entry><title>Unsupervised Data Augmentation for Consistency Training</title><id>https://techaiorbit.com/research/unsupervised-data-augmentation-for-consistency-training-190412848/</id><link href="https://techaiorbit.com/research/unsupervised-data-augmentation-for-consistency-training-190412848/"/><link rel="related" href="https://arxiv.org/abs/1904.12848"/><updated>2019-04-29T17:56:59.000Z</updated><summary>Semi-supervised learning lately has shown much promise in improving deep learning models when labeled data is scarce. Common among recent approaches is the use of consistency training on a large amount of unlabeled data to constrain model predictions to be invariant to input noise. In this work, we present a new perspective on how to effectively noise unlabeled examples and argue that the quality of noising, specifically those produced by advanced data augmentation methods, plays a crucial role in semi-supervised learning. By substituting simple noising operations with advanced data augmentation methods such as RandAugment and back-translation, our method brings substantial improvements across six language and three vision tasks under the same consistency training framework. On the IMDb text classification dataset, with only 20 labeled examples, our method achieves an error rate of 4.20, outperforming the state-of-the-art model trained on 25,000 labeled examples. On a standard semi-supervised learning benchmark, CIFAR-10, our method outperforms all previous approaches and achieves an error rate of 5.43 with only 250 examples. Our method also combines well with transfer learning, e.g., when finetuning from BERT, and yields improvements in high-data regime, such as ImageNet, whether when there is only 10% labeled data or when a full labeled set with 1.3M extra unlabeled examples is used. Code is available at https://github.com/google-research/uda.</summary></entry></feed>