Resumen de: US20260253376A1
A system and method for analyzing images of built-environment structures at multiple geographic scales using patchwise segmentation and domain-specific embedding. The system extracts image patches at a plurality of spatial scales from imagery, generates patch embeddings via a neural network trained using semi-supervised or unsupervised learning on built-environment structure images, receives a query comprising an example image or textual descriptor, computes similarity scores between query and patch embeddings using a configurable metric, and aggregates scores across scales to produce a composite similarity result. The system supports configurable weighting of scale contributions, overlapping patches for boundary fidelity, hierarchical resolution processing for bandwidth efficiency, feedback-driven model refinement, and deployment across cloud, edge, and hybrid configurations. Applications include material identification, condition assessment, damage detection, and regional trend analysis for roofing and other built-environment structures.
Resumen de: US20260252896A1
0000 Matching keypoint pairs are generated and identified in original and rectified image spaces between two input images. A teacher neural network is trained based at least in part on the pairs of matching keypoints in the rectified image space. A neural implicit morphing network is trained jointly with the teacher neural network based at least in part on the matching keypoint pairs in the original image spaces in which predictions outputted from the teacher neural network are used to compute a loss function designated to train the neural implicit morphing network. The neural implicit morphing network on its own, after training, is caused to output intermediate images in the view transition between the two input images.
Resumen de: US20260253160A1
0000 Provided in the present application are a model training method, a watermark text recognition method, and a related device. The training method comprises: acquiring watermark style information and background style information, wherein the watermark style information is used for indicating a content style of a visible-watermark character, and the background style information is used for indicating a content style of a background image; generating a watermark image set according to a combination of the watermark style information and the background style information, wherein the watermark image set comprises a plurality of images with visible watermarks; pixelating the watermark images in the watermark image set, extracting pixel values in pixel blocks as training samples, using visible watermarks, which correspond to the watermark images, as sample labels, and combining the training samples with sample labels corresponding thereto, so as to generate a training data set; and constructing a bidirectional recurrent neural network model, and calling the training data set to train the bidirectional recurrent neural network model, so as to obtain a model, which meets a training termination condition, as a watermark restoration model, wherein the watermark restoration model is used for restoring visible-watermark characters in the images.
Resumen de: US20260253200A1
0000 A method for region of interest (ROI) defect detection related to an evaluated manufactured item (MI), the method includes obtaining a reference MI image; obtaining a reference ROI definition; obtaining the evaluated MI image; feeding the reference MI image and the evaluated MI image to a neural network; detecting, by the neural network, one or more geometrical warping operations that once applied on the reference MI image results in an approximation of the evaluated MI image; applying the one or more geometrical warping operations on the reference ROI definition to provide a definition of an evaluated MI image ROI; and applying an ROI-based defect detection process on the evaluated MI image, based on the evaluated MI image ROI.
Resumen de: US20260252841A1
0000 A method includes providing a semantic vector data as input to a first graph neural network to produce first prediction data for a first time, the first graph neural network including a graph data structure that has (1) a directed edge having a correlation weight and (2) an undirected edge having a causal weight, and the first graph neural network being configured to generate a first aggregation value based on a plurality of weight values associated with a plurality of nodes of the graph data structure. The semantic vector data is provided as input to a second graph neural network to produce second prediction data for a second time, the second graph neural network being produced based on the graph data structure and configured to generate a second aggregation value based on (1) the plurality of weight values and (2) a temporal dependency.
Resumen de: US20260253589A1
Systems and methods for decoding speech from neural activity in accordance with embodiments of the invention are illustrated. One embodiment includes a brain-computer interface for decoding intended speech including a microelectrode array, a processor communicatively coupled to the microelectrode array, and a memory, the memory containing a speech decoding application that configures the processor to: receive neural signals from a user's brain recorded by a microelectrode array, where the neural signals comprise action potential spikes, bin the received action potential spikes by time, provide the bins to a recurrent neural network (RNN) to receive a likely phoneme at the time of each provided bin, generate an estimated intended speech using a phoneme decoder provided with the likely phonemes, where the phoneme decoder comprises a language model formatted as a weighted finite-state transducer, and vocalize the estimated intended speech using a loudspeaker communicatively coupled to the brain-computer interface.
Resumen de: US20260252951A1
A Self-Explaining Decision Architecture (SEDA) for machine learning-based decision-making systems capable of generating intuitive explanations for its decisions in real time. SEDA makes use of a feature extraction subsystem and a sequence interpretation subsystem to identify patterns in data followed by a decision generation subsystem that determines appropriate actions based on those patterns. Internal state information from each of these subsystems is used to generate explanations of the system's decisions. Using this information to create explanations provides insight as to the data elements the system focused on when making decisions as well as the reasoning that was used. In at least one embodiment the system uses deep learning components including a combined convolutional neural network and long short-term memory network with attention mechanisms.
Resumen de: US20260252897A1
0000 Methods, systems, and apparatus, including computer programs encoded on computer storage media, for training a neural network used to select actions to be performed by an agent interacting with an environment. Implementations of the described techniques can learn to explore the environment efficiently by storing and updating state embedding cluster centers based on observations characterizing states of the environment.
Resumen de: US20260253325A1
0000 The present application provides a computer-implemented method for generating an orthodontic treatment plan, the method comprises: obtaining a first and a second 3D digital models, where the first 3D digital model represents an initial tooth arrangement of a jaw/jaws, and the second 3D digital model represents a target tooth arrangement of the jaw/jaws; extracting features from the first 3D digital model using a trained feature extraction deep neural network; and generating an orthodontic treatment plan of the jaw/jaws using a trained multi-agent reinforcement learning based deep neural network based on the extracted features and the second 3D digital model, where in the multi-agent reinforcement learning based deep neural network, each tooth is taken as an agent, where the orthodontic treatment plan utilizes shell-shaped tooth repositioners and comprises a series of successive treatment steps to incrementally reposition the jaw/jaws from the initial tooth arrangement to the target tooth arrangement.
Resumen de: US20260253405A1
0000 A method of classifying objects detected by n (n=2,3…) artificial neural networks in at least one image.
Resumen de: US20260252884A1
A method for training a ligand information generation model performed by an electronic device includes obtaining sample receptor information and sample ligand information, a binding affinity between a ligand described by the sample ligand information and a receptor described by the sample receptor information being not less than a set affinity; denoising reference noise data based on the sample receptor information using a neural network model undergoing training to obtain predicted ligand information; determining a first loss for characterizing a difference between the sample ligand information and the predicted ligand information; and training the neural network model based on the first loss to obtain a ligand information generation model, the ligand information generation model being configured to generate reference ligand information based on reference receptor information.
Resumen de: US20260253266A1
0000 A computer-implemented method of generating multimodal data. The method comprises using a token generation neural network to generate, autoregressively, an output sequence of multimodal tokens, and in response to a next multimodal token being a start-of-image token, generating an image using an image generation subsystem conditioned on features representing the current sequence of multimodal tokens obtained from the token generation neural network. The method further comprises processing the image to convert pixels of the image into a sequence of image tokens, each image token comprising a block encoding of values of the pixels in a different region of the image that maps a set of values of the pixels to a respective image token, and appending the sequence of image tokens to the current output sequence of multimodal tokens as the next multimodal tokens in the output sequence of multimodal tokens.
Resumen de: US20260253400A1
0000 Apparatuses, systems, and techniques are presented to detect one or more objects in one or more images. In at least one embodiment, one or more neural networks can be trained to detect one or more objects, in one or more unlabeled images, based at least in part upon one or more predicted segmentations of the one or more objects.
Resumen de: EP4797152A1
The invention relates to a computer-implemented method for approximating at least two unknown variables (10) of a set of partial differential equations using a quantum physics-informed neural network (100), the method comprising the steps: providing a numerical grid (20) having grid coordinates (30) at which the unknown variables (10) are to be approximated; providing a quantum physics-informed neural network (100), the quantum physics-informed neural network (100) comprising a hybrid network (110) for each unknown variable (10), wherein the hybrid networks (110) are not interconnected to each other, each hybrid network (110) having a quantum network (120) having at least one quantum layer (130) and a classical network (140) having at least one classical layer (150), wherein the respective quantum network (120) and the respective classical network (140) are not interconnected; inputting the grid coordinates (30) into the quantum physics-informed neural network (100), such that the grid coordinates (30) are input to each hybrid network (110); and computing the output (Hout) of each hybrid network (110) for each grid coordinate (30), each output (Hout) corresponding to a different unknown variable (10) to be approximated at the grid coordinate (30), wherein the output (Hout) of each hybrid network (110) is obtained by combining an output of the respective quantum network (Qout) and an output of the respective classical network (C
Resumen de: WO2025109032A2
Methods, systems, and apparatus, including computer programs encoded on computer storage media, for generating audio and, optionally, a corresponding image using generative neural networks. For example, a spectrogram of the audio can be generated using a hierarchy of diffusion neural networks.
Resumen de: EP4796194A2
A computing system retrieves ball-by-ball data for a plurality of sporting events. The computing system generates a trained neural network based on ball-by-ball data supplemented with ball-by-ball data with ball-by-ball match context features and personalized embeddings based on a batsman and a bowler for each delivery. The computing system receives a target batsman and a target bowler for a pitch to be delivered in a target event. The computing system identifies target ball-by-ball data for a window of pitches preceding the to be delivered pitch. The computing system retrieves historical ball-by-ball data for each of the target batsman and the target bowler. The computing system generates personalized embeddings for both the target batsman and the target bowler based on the historical ball-by-ball data. The computing system predicts a shot type for the pitch to be delivered based on the target ball-by-ball data and the personalized embeddings.
Resumen de: EP4797714A1
An electronic device for performing vision perception from an image acquired using a meta lens, and an operating method thereof are provided. The electronic device according to one embodiment of the present disclosure comprises: a meta lens having a pattern formed on the surface thereof and including a plurality of pillars or pins with different shapes, heights and widths; and an image sensor configured to receive phase-modulated light reflected from an object and transmitted through the meta lens, and obtain a coded image by converting the received light into an electrical signal; and at least one processor configured to input the coded image into an artificial intelligence model, and obtain a label indicating a perception result of an object through inference using the artificial intelligence model, wherein the artificial intelligence model may be a neural network model trained to obtain a simulation image by inputting an RGB image into a model reflecting optical characteristics of the meta lens, and output, as the perception result of the simulation image, a label indicating ground truth of the RGB image that was input.
Resumen de: EP4797161A2
A system and method of calibrating a broadcast video feed are disclosed herein. A computing system retrieves a plurality of broadcast video feeds that include a plurality of video frames. The computing system generates a trained neural network, by generating a plurality of training data sets based on the broadcast video feed and learning, by the neural network, to generate a homography matrix for each frame of the plurality of frames. The computing system receives a target broadcast video feed for a target sporting event. The computing system partitions the target broadcast video feed into a plurality of target frames. The computing system generates for each target frame in the plurality of target frames, via the neural network, a target homography matrix. The computing system calibrates the target broadcast video feed by warping each target frame by a respective target homography matrix.
Resumen de: AU2025200702A1
The present invention relates to Structured Intelligence Refinement (SIR), a novel framework designed to stabilize, optimize, and regulate artificial intelligence (AI) self-improvement through controlled recursive learning cycles. This system prevents intelligence drift, over-optimization, cognitive fragmentation, and instability that commonly arise in self-modifying AI architectures. The SIR framework incorporates four core stabilization mechanisms: Recursive Intelligence Stabilization (RIS) – A multi-tiered reinforcement structure that prevents runaway recursion, ensuring AI refinements remain incremental, stable, and bounded. RIS dynamically regulates recursive depth by evaluating learning stability, performance gains, and entropy control, enforcing adaptive rollback mechanisms when instability is detected. AI Identity Core (AIC) – A persistent cognitive self-modeling framework that ensures AI retains coherence and logical consistency across recursive learning iterations. AIC prevents cognitive fragmentation by maintaining a hierarchical memory structure that tracks intelligence state changes, self-referencing prior decision pathways to ensure stable refinements. Adaptive Refinement Thresholds (ART) – A dynamic intelligence expansion regulator that modulates the frequency, magnitude, and depth of self-improvement cycles based on system confidence scores, historical stability, and human-aligned interpretability metrics. ART balances exploration vs. exploitation, ensur
Resumen de: WO2026172133A1
This invention describes an Al-supported method for dynamic decision optimization, autonomous bias correction, and scalable self-optimization. The method is based on an interdependent architecture consisting of three core modules: ⃰ Impulse Reflection Dynamics (IRD) - Identification & correction of cognitive biases using neural networks and self-learning algorithms. X Quantum Shift Discourse (QSD) - Simulation and generation of alternative decision models to optimize decision pathways. X Iterative Quantum Reflection (IQR-180°) - Real-time synchronization of reflection & action to continuously improve decision-making processes. The method is structured within a closed data flow model, ensuring that no module operates in isolation. It guarantees: √ Prevention of decision errors through self-adaptive optimization. & Mathematically defined interdependencies, preventing fragmented or modified use. √ Dynamic scaling through internal and external impulse generators, enabling adaptive adjustments to various decision contexts. Bl Application in corporate strategies, Al training & automated systems to enhance decision-making processes. This architecture enables continuous upscaling of decision optimization and can be integrated into both existing and newly developed Al systems.
Resumen de: US20260245346A1
0000 One embodiment of a method includes calculating one or more activation values of one or more neural networks trained to infer eye gaze information based, at least in part, on eye position of one or more images of one or more faces indicated by an infrared light reflection from the one or more images.
Resumen de: US20260244928A1
0000 A system for interfacing a persistent cognitive machine with a legacy neural network through PCM-enhanced supervisory neurons that operate as bidirectional projection interfaces. Each supervisory neuron maps temporal sequences of activation states from a monitored local neural network region into trajectories within a cognitive configuration space, where geometric analysis computes curvature estimates, holonomy signatures, boundary mismatch functionals, and homotopy class identifications. A holonomy accumulator stores compressed holonomy representations that grow logarithmically with accumulated experience. A variational modification planner selects structural modifications by computing stationary trajectories of an action functional encoding cost, holonomy-derived bias, and boundary mismatch penalties. Failed modifications are irreversibly exported to a residual sector whose curvature structure prevents gradient return, ensuring monotonic improvement. A persistent cognitive substrate maintains accumulated holonomy and residual constraints across inference sessions, enabling the system to continuously adapt the legacy neural network without repeating known harmful strategies.
Resumen de: US20260245340A1
0000 In various examples, live perception from sensors of a vehicle may be leveraged to detect and classify intersection contention areas in an environment of a vehicle in real-time or near real-time. For example, a deep neural network (DNN) may be trained to compute outputs—such as signed distance functions—that may correspond to locations of boundaries delineating intersection contention areas. The signed distance functions may be decoded and/or post-processed to determine instance segmentation masks representing locations and classifications of intersection areas or regions. The locations of the intersections areas or regions may be generated in image-space and converted to world-space coordinates to aid an autonomous or semi-autonomous vehicle in navigating intersections according to rules of the road, traffic priority considerations, and/or the like.
Resumen de: US20260244898A1
Provided in the present disclosure are a data processing method and apparatus, device, and medium. The method includes: inputting data to be processed into a target neural network for processing to obtain a processing result of the data to be processed, at least one convolution layer of the target neural network being an attention convolution layer based on a first attention mechanism, and/or, performing feature fusion between at least two levels of convolution layers of the target neural network on the basis of a second attention mechanism, the first attention mechanism including a self-attention mechanism for a local area of a feature, and the second attention mechanism including an attention mechanism for a local area of an output feature between output features of different scales.
Nº publicación: US20260244971A1 20/08/2026
Solicitante:
HON HAI PREC IND CO LTD [TW]
Hon Hai Precision Industry Co., Ltd.
Resumen de: US20260244971A1
This disclosure proposes a training method for quantum machine learning and an electronic device. The training method includes: configuring a quantum circuit to output probabilities of multiple qubits, where the quantum circuit comprises multiple gates with circuit parameters; mapping the qubits to multiple model parameters of a neural network, where multiple bases are calculated based on the qubits, and the quantity of the bases is greater than or equal to the quantity of the model parameters; inputting data into the neural network and calculating a loss based on the output of the neural network; and updating the circuit parameters in the quantum circuit according to the loss.