RESEARCH ARTICLES
Manner of death, causes of death and autopsies in infants, children and adolescents: An overview from a German metropolis 2002–2012
Nov 2021 – International Journal of Legal Medicine (Springer Nature)
DOI: 10.1007/s00194-022-00568-y
Fractures and skin lesions in pediatric abusive head trauma: a forensic multi-center study
Dec 2021 – International Journal of Legal Medicine (Springer Nature)
DOI: 10.1007/s00414-021-02751-4
Abusive head trauma in court: a multi-center study on criminal proceedings in Germany
Oct 2020 – International Journal of Legal Medicine (Springer Nature)
DOI: 10.1007/s00414-020-02435-5
Post-mortem estimation of gestational age and maturation of new-borns by CT examination of clavicle length, femoral length and femoral bone nuclei
Jun 2020 – Forensic Science International [ACKNOWLEDGED SUPPORT]
DOI: 10.1016/j.forsciint.2020.110391
Pure Functions in C: A Small Keyword for Automatic Parallelization
May 2020 – International Journal of Parallel Programming
DOI: 10.1007/s10766-020-00660-4
Extending PluTo for Multiple Devices by Integrating OpenACC
March 2018 – 26th Euromicro International Conference on Parallel, Distributed and Network-based Processing (PDP)
DOI: 10.1109/PDP2018.2018.00049
Pure Functions in C: A Small Keyword for Automatic Parallelization
September 2017 – IEEE International Conference on Cluster Computing (CLUSTER)
DOI: 10.1109/CLUSTER.2017.32
Energy-Efficiency and Performance Comparison of Aerosol Optical Depth Retrieval on Distributed Embedded SoC Architectures
July 2017 – In book: Scientific Computing and Algorithms in Industrial Simulations (pp.341-358)
DOI: 10.1007/978-3-319-62458-7_17
VarySched: A Framework for Variable Scheduling in Heterogeneous Environments
September 2016 – IEEE International Conference on Cluster Computing (CLUSTER)
DOI: 10.1109/CLUSTER.2016.19
An efficient geosciences workflow on multi-core processors and GPUs: a case study for aerosol optical depth retrieval from MODIS satellite data
February 2016 – International Journal of Digital Earth 9(8):1-18
DOI: 10.1080/17538947.2015.1130087
Comparison of Acceleration Techniques for Selected Low-Level Bioinformatics Operations + Supplementary Material
February 2016 – Frontiers in Genetics 7
DOI: 10.3389/fgene.2016.00005
Impact of the Scheduling Strategy in Heterogeneous Systems That Provide Co-Scheduling
January 2016 – 1st COSH Workshop on Co-Scheduling of HPC Applications
DOI: 10.14459/2016md1286954
Multicore Processors and Graphics Processing Unit Accelerators for Parallel Retrieval of Aerosol Optical Depth From Satellite Data: Implementation, Performance, and Energy Efficiency
June 2015 – IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 8(5):2306-2317
DOI: 10.1109/JSTARS.2015.2438893
Hardware-Aware Automatic Code-Transformation to Support Compilers in Exploiting the Multi-Level Parallel Potential of Modern CPUs
February 2015 – COSMIC ’15 Proceedings of the 2015 International Workshop on Code Optimisation for Multi and Many Cores
DOI: 10.1145/2723772.2723776
Facilitate SIMD-Code-Generation in the Polyhedral Model by Hardware-aware Automatic Code-Transformation
January 2013 – IMPACT 2013 Volume: 45
DOI: 10.13140/2.1.5066.3368
PATENTS
Feld et al.
Method and computer program for determining a placement of at least one circuit for a reconfigurable logic device
Verfahren und Computerprogramm zur Bestimmung einer Positionierung von mindestens einer Schaltung für eine rekonfigurierbare logische Vorrichtung
Embodiments relate to a method and computer program for determining a placement of at least one circuit for a reconfigurable logic device. The method comprises obtaining (110) information related to the at least one circuit. The at least one circuit comprises a plurality of blocks and a plurality of connections between the plurality of blocks. The plurality of blocks comprise a plurality of logic blocks. The method further comprises calculating (120) a circuit graph based on the information related to the at least one circuit. The circuit graph comprises a plurality of nodes and a plurality of edges. The plurality of nodes represent at least a subset of the plurality of blocks of the at least one circuit and wherein the plurality of edges represent at least a subset of the plurality of connections between the plurality of blocks of the at least one circuit. The method further comprises determining (130) a force-directed layout of the circuit graph. The force-directed layout is based on attractive forces based on the plurality of connections between the plurality of blocks and based on repulsive forces between the plurality of blocks. The method further comprises determining (140) a placement of the plurality of logic blocks onto a plurality of available logic cells of the reconfigurable logic device based on the force-directed layout of the circuit graph.
LISTING IN THE EUROPEAN PATENT OFFICE
Application number: EP20160203521 20161212
Priority number(s): EP20160203521 20161212
Also published as:
CN108228972 (B)
CN108228972 (A)
US10460062 (B2)
US2018165400 (A1)
THESIS‘
FieldPlacer – A flexible, fast and unconstrained force-directed placement method for heterogeneous reconfigurable logic architectures. Dissertation, Universität zu Köln.
The field of placement methods for components of integrated circuits, especially in the domain of reconfigurable chip architectures, is mainly dominated by a handful of concepts. While some of these are easy to apply but difficult to adapt to new situations, others are more flexible but rather complex to realize. This work presents the FieldPlacer framework, a flexible, fast and unconstrained force-directed placement method for heterogeneous reconfigurable logic architectures, in particular for the ever important heterogeneous FPGAs. In contrast to many other force-directed placers, this approach is called ‘unconstrained’ as it does not require a priori fixed logic elements in order to calculate a force equilibrium as the solution to a system of equations. Instead, it is based on a free spring embedder simulation of a graph representation which includes all logic block types of a design simultaneously. The FieldPlacer framework offers a huge amount of flexibility in applying different distance norms (e. g., the Manhattan distance) for the force-directed layout and aims at creating adapted layouts for various objective functions, e. g., highest performance or improved routability. Depending on the individual situation, a runtime-quality trade-off can be considered to either produce a decent placement in a very short time or to generate an exceptionally good placement, which takes longer. An extensive comparison with the latest simulated annealing placement method from the well-known Versatile Place and Route (VPR) framework shows that the FieldPlacer approach can create placements of comparable quality much faster than VPR or, alternatively, generate better placements in the same time. The flexibility in defining arbitrary objective functions and the intuitive adaptability of the method, which, among others, includes different concepts from the field of graph drawing, should facilitate further developments with this framework, e. g., for new upcoming optimization targets like the energy consumption of an implemented design.
Effiziente Vektorisierung durch semi-automatisierte Code-Optimierung im Polyedermodell
In den vergangenen zwei Jahrzehnten haben sich Vektoreinheiten als Beschleuniger in CPUs auch im Bereich der PCs etabliert. Allerdings sind selbst moderne Compiler nicht generell in der Lage, dieses Potenzial auszuschöpfen, wenn der Programmierer die entsprechenden Codes nicht explizit für den jeweiligen Beschleuniger schreibt. Transformationen des Codes, die für eine Nutzung der Vektoreinheiten nötig wären, können von vielen aktuellen Compilern nicht oder nicht immer realisiert werden, da die hierzu nötigen mathematischen Operationen und Analysen nicht implementiert oder noch gar nicht entwickelt sind. In dieser Arbeit wurden Methoden zur Transformation von Quellcodes entwickelt und implementiert, die auf Gomory-Cut-Lösungsverfahren der ganzzahligen Optimierung und Konzepten der Lineare n Algebra beruhen und als Vorstufe einer End-Kompilierung (durch Compiler wie den GCC oder ICC) eingesetzt werden können. Der Fokus lag auf der Entwicklung von Codeoptimierungen für eine Vektorisierung von Schleifen in C-Codes durch ein semi-automatisches Framework. Ziel der Transformationen ist eine bessere Nutzung von SIMD-Architekturen, z. B. den SSE-Einheiten aktueller x86-Prozessoren, die Prinzipien lassen sich aber auch auf andere Hardware-Architekturen mit vektorbasierten Befehlssätzen übertragen. Neben Transformationen zur Ermöglichung der Vektorisierung steht die automatische Anpassung an die vorliegende Hardware im Fokus der Betrachtung um eine für die Speichernutzung optimale Transformation durchzuführen. Durch die entwickelten Methoden können erhebliche Speedups erreicht werd en, die durch die Speicherzugriffsoptimierung z.T. deutlich über dem durch die parallelen Berechnungen möglichen Speedup liegen.
Neueste Kommentare